March 2023. GPT-4 launches at $30 per million input tokens and $60 per million output.
July 2026. GPT-5.6 Luna: $0.20 in, $1.20 out. The median across 143 tracked models sits at $1 and $4.
That is the steepest sustained price collapse in the history of enterprise software. Roughly 99% off the input price of frontier-grade intelligence in three and a half years.
Now the other number. Eric Glyman runs Ramp, a company whose entire product is watching where corporate money goes. Here is what he found when he looked at his own.
Glyman
AI is extremely good at spending your money very quietly. Our own token spend went from a rounding error to more than 10% of payroll in a year. One week in May we burned through $1.5m. Our CFO didn't love telling me that number, and he really didn't love telling the internet.
Ramp decided that was the fastest-growing line in the business and shipped a product to help other companies track theirs. Over a thousand signed up.
They are not an outlier. Sahil Lavingia reported that June 2026 was the first month Gumroad spent equal amounts on AI tokens and human payroll.
Prices fell 99%. Bills went up. Both true, same three years.
The chart everyone draws wrong
Plotted as one line, this looks like a simple story about things getting cheaper. Plotted honestly it is two lines, and the gap between them is the entire point.
The floor collapsed. Cheap capable intelligence is now effectively free at the margin.
The frontier did not. Adjusted for what they can do, Fable 5 and Opus 5 are extraordinary value. In absolute dollars per token, the top of the market costs roughly what GPT-4 cost at launch.
Amended two days after publishing. On August 12, xAI shipped Grok 4.6 at $2 in and $6 out, scoring 61 on the Artificial Analysis index against Fable 5's 62 and Opus 5's 63. A model one point off the top of the board, at a fifth of the price. That is not the floor falling. That is the frontier getting cheap, which is the thing this section says is not happening.
The conclusion below holds: routing beats rates. The framing above needs the amendment, and the full argument is here.
The middle is filling in too. On August 10 Anthropic made Claude Sonnet 5's introductory pricing permanent at $2 in and $10 out, the rate it launched with in June. A capable mid-tier at a fifth of Opus 5 and a fiftieth of Fable's output rate is exactly the kind of option that makes the routing argument below concrete rather than theoretical.
So nothing got cheaper for anyone doing frontier work. What happened is that a floor appeared underneath, and the interesting question stopped being "what does a token cost" and became "which line is this workload running on."
Why the bill grew anyway
Jevons, obviously. Cheaper unit, more units. But the specific mechanism is worth naming, because it is new and it changes how you budget.
A chat turn is one call. You control it. An agent running for four hours is thousands of calls, and the agent decides how many. You did not authorise 40,000 tokens of exploration. You authorised a goal. Consumption moved from something you approve to something you discover afterwards.
Cursor published the cleanest illustration anyone has produced. They pointed a team of agents at SQLite's 835-page manual and asked for a reimplementation in Rust. It worked. The result passed 100% of a held-out test suite. And the cost of getting there varied 15x depending on which mix of models ran the job.
Same task. Same passing result. Fifteen times the spread. That is not a pricing problem. That is an allocation problem wearing a pricing problem's clothes.
The lever is routing, not rates
Everyone's instinct is to go and negotiate. The actual lever is deciding which model touches which step.
The pattern that consolidated across the market this quarter: a frontier model as orchestrator, cheap models as workhorses. The expensive thing does planning, decomposition, and judgment on the hard call. Everything downstream of that runs near the floor. Cursor productised it in July with Router, claiming frontier-quality output at 60% lower cost by choosing the model per request rather than per user.
That is the same move Coinbase made when it halved its bill without capping anyone, which was the finding at the centre of Part 1 of this series. Nothing since has changed the conclusion. Capping spend slows the company down. Allocating spend does not.
The tooling is catching up to the idea. Claude Managed Agents shipped per-session budgets this month: set a limit, and a session that reaches it pauses and emits a budget_reached event rather than quietly continuing. Cost control is moving into the runtime, where it belongs, instead of living in a monthly invoice review.
What to actually change
Stop measuring cost per token. It is the wrong denominator, and it is falling so fast that it flatters you. Measure cost per completed task. Cost per closed ticket, per reviewed contract, per qualified lead. That is the number that tells you whether the workflow is worth running, and the only one a CFO can compare against the alternative, which is a person.
Instrument by workflow, not by seat. Seat-level spend tells you who logged in. Workflow-level spend tells you that your support triage loop costs $0.40 a ticket and your contract review loop costs $19, and that one of those needs looking at.
Put a budget in the runtime. If your agents can spend without a ceiling that pauses them, you do not have a budget. You have a forecast.
Assume the floor keeps falling. Anything you are not doing today because it is too expensive per token is worth re-costing every quarter. The workloads that were uneconomic in March were economic by July.
Token prices are not your problem. Nobody in this market is losing on rates. They are losing by running frontier intelligence over work that a model costing 1% as much would have finished correctly.
Route it. Don't cap it.
Route it. Don’t cap it.
Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We take one live workflow, work out what it actually costs per completed task, and where a cheaper model would have finished the job.
Book the Diagnostic →budget_reached event.