LiteInk LiteInk

The Jevons Paradox in AI: Why Cheaper Tokens Are Burning Bigger Budgets

As inference costs plummet, AI usage is exploding. The economic principle from 1865 explains why your AI bill keeps going up, not down.

The Jevons Paradox in AI: Why Cheaper Tokens Are Burning Bigger Budgets illustration

In 2025, the blended cost of AI tokens fell 67% year-over-year — from $18.40 to $6.07 per million tokens. In that same period, Uber exhausted its entire 2026 AI coding budget by April. Microsoft quietly revoked its developers’ Claude Code licenses. And one unnamed company discovered a $500 million Claude bill on its invoice because nobody had bothered to set usage limits.

Token prices are cratering. Token bills are exploding. A 19th-century economist could have predicted this.


In 1865, William Stanley Jevons wrote a book about coal. Steam engines were getting more efficient. You’d think that means less coal burned. The opposite happened — cheaper, more efficient engines made coal viable for more industries. Total consumption went up. His name got attached to the pattern: the Jevons Paradox.

It’s been observed in electricity, gasoline, bandwidth. And now it’s happening to AI tokens.

What’s actually happening out there

Uber capped employee AI spending after blowing through its full 2026 AI budget in four months. Microsoft pulled Claude Code licenses from developers months after rolling them out. A Priceline employee told TechCrunch that a routine Cursor contract renewal came back four to five times more expensive than the previous term.

Here’s the number that should make every CFO uncomfortable: enterprises waste an estimated 26% of all AI spend, according to a Harness survey. One in five companies now spend over $1 million per month on AI. That’s $260,000 evaporating every thirty days with no clear ownership.

Gartner puts worldwide AI spending at $2.59 trillion in 2026, up 47%. Generative AI model spend more than doubled. Meanwhile, analysis of 2.4 billion enterprise API calls shows the per-token cost dropped 67%.

Cheaper units. Higher total spend. The Jevons Paradox requires two conditions: demand must respond strongly to price, and lower prices must unlock new use cases that weren’t economical before. AI tokens meet both.

Agents changed the math

In late 2025, Anthropic shipped Claude Opus 4.5, OpenAI released GPT-5.1, Google launched Gemini 3 Pro. These were the first models good enough to power genuinely autonomous agents — tools that don’t just suggest code or text but execute multi-step tasks, calling models hundreds or thousands of times per workflow.

Agentic features multiplied per-developer token consumption by 18.6x in nine months, according to Jellyfish, an engineering management platform. Their two-year study of 20,000 developers found engineers who used AI tools the most were roughly twice as productive — but consumed 10x the tokens to get there.

Vitaly Gordon, CEO of Faros AI, relayed a conversation that captures the executive dilemma: “One of my engineers spent $40,000 on tokens last month, and I genuinely don’t know whether I should stop him or tell everyone else to be like him.”

That’s the problem. Organizations can’t tell productive spend from waste. The 26% waste figure exists because most companies can measure spending but not returns. They can’t measure token-level ROI. A bill arrives. Someone pays it. Nobody knows if it produced proportional value.

This has happened before

Two times, actually.

The first was telecom in the late 1990s. Companies had no visibility into long-distance and international call costs until bills arrived. An entire industry — telecom expense management — emerged to audit and optimize. The second was cloud. AWS launched in 2006, and within five years companies were bleeding money on idle instances and forgotten staging environments. The FinOps Foundation formed under the Linux Foundation to bring discipline to cloud spend. It’s now standard practice.

AI tokens are the third wave. The scale is different, though. J.R. Storment, executive director of the FinOps Foundation, told TechCrunch: “Tracking cloud costs is a hundreds-of-millions-of-rows-a-month data problem. Tracking token costs is a trillions-of-rows-a-month data problem.”

The Linux Foundation has launched the Tokenomics Foundation to build open standards for AI token usage — metrics like cost-per-intelligence and tokens-per-watt. Startups like Pay-i, Paid, and Factory are building tools to track and route token spend. Ramp, Datadog, New Relic are bolting token observability onto existing platforms.

Tooling won’t solve the fundamental tension, though. Priceline’s senior director of IT finance, Chris Reed, compared the dynamic to a drug epidemic: “They let you try it to get you hooked on it, and now you’re kind of beholden to it.” Priceline has started placing token limits on certain teams — treating AI spending as a behavioral problem, not just a technical one.

What to watch

Every lab — OpenAI, Anthropic, Google, Moonshot, xAI — is simultaneously dropping prices and shipping more capable agentic features that consume exponentially more tokens. Cheaper models will accelerate consumption, not contain it.

Three things will determine who navigates this well.

First: whether the industry adopts a shared vocabulary for token economics. The Tokenomics Foundation is promising standards by late 2026. Without them, every vendor defines “usage” differently. Priceline is already finding discrepancies between vendor-reported usage and internal data.

Second: model routing. Factory launched a router that automatically picks the cheapest model capable of handling each task. Anthropic already silently routes Claude requests to cheaper Sonnet or Haiku models when full Opus power isn’t needed. This will become default, not a premium feature.

Third: connecting token spend to business outcomes. Nicholas Arcolano, head of research at Jellyfish: “Whether extreme spend pays off comes down to the ultimate business value of shipped code, which most companies still can’t measure.” The companies that figure out that measurement layer first will have an enormous cost advantage.

For builders and indie developers: if the largest enterprises on earth — with dedicated finance teams and procurement departments — can’t control their AI bills, the risk is proportionally higher for smaller teams. The era of “let’s just enable AI for everyone and see what happens” is ending. Tokens are becoming a variable cost that needs to be budgeted, monitored, and tied to outcomes.

The tokens got cheaper. The discipline got more expensive.


Disclosure: This is independent analysis, not investment or professional advice.

ESC