The Cheaper It Gets, the More You'll Spend
The price war everybody just cheered is the thing that's going to blow up your AI budget. Not despite the falling prices — because of them. Cheaper tokens don't shrink an AI bill. They manufacture a bigger one. And the providers underpricing each other into the floor this month understand that better than the customers celebrating the discount.
Look at what actually happened in the last few weeks. OpenAI opened all three GPT-5.6 tiers to everyone on July 9 — Luna landing frontier-family capability at a dollar per million input tokens, six on output. Grok 4.5 shipped in early July pitched explicitly as Opus-class but faster, leaner, cheaper. Gemini 3.6 Flash dropped on the 21st. Every headline reads the same way: inference is collapsing toward free, the bottleneck is lifting, go build. It scans like a discount season. It's the setup for the opposite.
Jevons called this in 1865
The mechanism is a hundred and sixty years old and it has a name. William Stanley Jevons noticed that when James Watt's engine made steam power dramatically more efficient with coal, England didn't burn less coal. It burned vastly more. Efficiency didn't reduce consumption — it made coal economical for a hundred uses that weren't worth it before, and total demand exploded past the per-unit saving. Cheaper made the resource worth reaching for in places nobody would have bothered. That's the whole trap, and inference is walking straight into it.
Here are the two numbers that matter, and the space between them is your invoice. Token prices for a given tier of capability have fallen something like 98% since the early GPT-3 days — roughly twenty dollars per million tokens in late 2022 down to well under fifty cents now. Over the same stretch, the average enterprise AI budget went from about $1.2 million a year in 2024 to around $7 million in 2026. Prices cratered. Spend sextupled. Those aren't contradictory readings of the market. They're the same fact seen from two ends.
The furnace nobody priced in
Here's the part that's genuinely wild, the part I keep turning over. What made spend detonate isn't people running their old workloads a little more. It's an entirely new shape of workload that only became affordable once tokens got cheap: the agent.
A single-turn query is one pass. Ask, answer, done — a bounded, predictable sip. An agent doesn't sip. It loops. It calls a tool, reads the result, decides to call another, chains a dozen reasoning steps to run down one lead, backtracks, retries. The measured cost of an agentic task runs anywhere from 100 to 1,000 times the tokens of the single-turn version of the same job. And agents only became economically viable because the per-token price fell far enough to make that thrashing affordable. Read that as the loop it is: the price drop didn't save anyone money. It unlocked the single most token-hungry pattern in computing and handed everyone permission to run it. You didn't get a discount. You got a green light to build the expensive thing you couldn't justify last quarter.
The scale of it barely fits in the head. Google was processing something like 9.7 trillion tokens a month in mid-2024. A year later, 480 trillion. This past May, 3,200 trillion — a sevenfold jump in a single year. Enterprise consumption across the board rose more than 100x while the per-token price fell only about 10x. Multiply a 10x price cut against a 100x volume surge and you don't get relief. You get a bill an order of magnitude bigger than the one the cheaper tokens were supposed to spare you.
The flagship never actually dropped
Now the tell, the detail the price-war coverage skates right past. The flagship didn't get cheaper. GPT-5.6's top tier held at five dollars in, thirty out — identical to what the previous generation charged at launch in April, a full generation of capability gains absorbed with the sticker price frozen. The reasoning premium — what you pay for the model to actually think rather than autocomplete — still runs better than 30x the non-reasoning rate. So the "war" is a fire sale on the tier you'll use least, the cheap non-reasoning one, while the tier that does your genuinely hard work holds firm. The number in the headline and the number on your statement are not the same number. One is marketing. The other is the reasoning tokens your agents actually burn.
The fair pushback: falling unit costs are real, and for a fixed workload they'd genuinely save money. Correct — and that's exactly the assumption that never holds. Jevons only bites because demand is elastic. If you were going to run the same ten queries regardless of price, cheaper tokens are a straight win. Nobody runs the same ten queries. The instant the price drops, the ten thousand queries that were uneconomical yesterday become obvious today, and you run those instead. Elastic demand is the entire ballgame, and enterprise AI demand is about as elastic as anything in the history of computing.
Run the inversion Munger would run. Don't ask what a token costs. Ask what a falling token price unlocks that you'll feel obligated to build — and budget that, because that's the line that's going to the moon while the unit price goes to zero. Both are true at once. Put a governor on consumption before the price war convinces your org it doesn't need one.
The tokens are getting cheaper. Your bill is not. Anyone selling you the first fact and leaving out the second is selling you the coal.
The meter reads pennies a token right now. That's not the good news. That's the invitation, and it's addressed to your budget.
— Dustin