March 2023. GPT-4 launches at $30 per million input tokens and $60 per million output.

July 2026. GPT-5.6 Luna: $0.20 in, $1.20 out. The median across 143 tracked models sits at $1 and $4.

That is the steepest sustained price collapse in the history of enterprise software. Roughly 99% off the input price of frontier-grade intelligence in three and a half years.

Now the other number. Eric Glyman runs Ramp, a company whose entire product is watching where corporate money goes. Here is what he found when he looked at his own.

Eric Glyman
Eric
Glyman

AI is extremely good at spending your money very quietly. Our own token spend went from a rounding error to more than 10% of payroll in a year. One week in May we burned through $1.5m. Our CFO didn't love telling me that number, and he really didn't love telling the internet.

Eric Glyman, cofounder and CEO of Ramp · July 2026

Ramp decided that was the fastest-growing line in the business and shipped a product to help other companies track theirs. Over a thousand signed up.

They are not an outlier. Sahil Lavingia reported that June 2026 was the first month Gumroad spent equal amounts on AI tokens and human payroll.

Prices fell 99%. Bills went up. Both true, same three years.

The chart everyone draws wrong

Plotted as one line, this looks like a simple story about things getting cheaper. Plotted honestly it is two lines, and the gap between them is the entire point.

Fig. 1
The floor collapsed. The frontier did not.
Two lines, not one USD PER MILLION INPUT TOKENS LOG SCALE · MAR 2023 TO JUL 2026 $50 $10 $1 $0.10 GPT-4 $30 TURBO $10 4o $2.50 FABLE 5 $10 OPUS 5 $5 NANO $0.10 LUNA $0.20 MAR 23 NOV 23 MAY 24 APR 25 JUL 26 FRONTIER FLAGSHIP CHEAPEST CAPABLE MEDIAN OF 143 MODELS, AUG 2026: $1 IN / $4 OUT
The frontier costs roughly what GPT-4 cost at launch. What appeared underneath it is a floor 50 times lower.
Sources: OpenAI and Anthropic published API pricing, March 2023 to July 2026; BenchLM median across 143 models, August 6, 2026

The floor collapsed. Cheap capable intelligence is now effectively free at the margin.

The frontier did not. Adjusted for what they can do, Fable 5 and Opus 5 are extraordinary value. In absolute dollars per token, the top of the market costs roughly what GPT-4 cost at launch.

Amended two days after publishing. On August 12, xAI shipped Grok 4.6 at $2 in and $6 out, scoring 61 on the Artificial Analysis index against Fable 5's 62 and Opus 5's 63. A model one point off the top of the board, at a fifth of the price. That is not the floor falling. That is the frontier getting cheap, which is the thing this section says is not happening.

The conclusion below holds: routing beats rates. The framing above needs the amendment, and the full argument is here.

The middle is filling in too. On August 10 Anthropic made Claude Sonnet 5's introductory pricing permanent at $2 in and $10 out, the rate it launched with in June. A capable mid-tier at a fifth of Opus 5 and a fiftieth of Fable's output rate is exactly the kind of option that makes the routing argument below concrete rather than theoretical.

So nothing got cheaper for anyone doing frontier work. What happened is that a floor appeared underneath, and the interesting question stopped being "what does a token cost" and became "which line is this workload running on."

Why the bill grew anyway

Jevons, obviously. Cheaper unit, more units. But the specific mechanism is worth naming, because it is new and it changes how you budget.

A chat turn is one call. You control it. An agent running for four hours is thousands of calls, and the agent decides how many. You did not authorise 40,000 tokens of exploration. You authorised a goal. Consumption moved from something you approve to something you discover afterwards.

Cursor published the cleanest illustration anyone has produced. They pointed a team of agents at SQLite's 835-page manual and asked for a reimplementation in Rust. It worked. The result passed 100% of a held-out test suite. And the cost of getting there varied 15x depending on which mix of models ran the job.

Same task. Same passing result. Fifteen times the spread. That is not a pricing problem. That is an allocation problem wearing a pricing problem's clothes.

The lever is routing, not rates

Everyone's instinct is to go and negotiate. The actual lever is deciding which model touches which step.

The pattern that consolidated across the market this quarter: a frontier model as orchestrator, cheap models as workhorses. The expensive thing does planning, decomposition, and judgment on the hard call. Everything downstream of that runs near the floor. Cursor productised it in July with Router, claiming frontier-quality output at 60% lower cost by choosing the model per request rather than per user.

That is the same move Coinbase made when it halved its bill without capping anyone, which was the finding at the centre of Part 1 of this series. Nothing since has changed the conclusion. Capping spend slows the company down. Allocating spend does not.

The tooling is catching up to the idea. Claude Managed Agents shipped per-session budgets this month: set a limit, and a session that reaches it pauses and emits a budget_reached event rather than quietly continuing. Cost control is moving into the runtime, where it belongs, instead of living in a monthly invoice review.

What to actually change

Stop measuring cost per token. It is the wrong denominator, and it is falling so fast that it flatters you. Measure cost per completed task. Cost per closed ticket, per reviewed contract, per qualified lead. That is the number that tells you whether the workflow is worth running, and the only one a CFO can compare against the alternative, which is a person.

Instrument by workflow, not by seat. Seat-level spend tells you who logged in. Workflow-level spend tells you that your support triage loop costs $0.40 a ticket and your contract review loop costs $19, and that one of those needs looking at.

Put a budget in the runtime. If your agents can spend without a ceiling that pauses them, you do not have a budget. You have a forecast.

Assume the floor keeps falling. Anything you are not doing today because it is too expensive per token is worth re-costing every quarter. The workloads that were uneconomic in March were economic by July.

Token prices are not your problem. Nobody in this market is losing on rates. They are losing by running frontier intelligence over work that a model costing 1% as much would have finished correctly.

Route it. Don't cap it.

Route it. Don’t cap it.

Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We take one live workflow, work out what it actually costs per completed task, and where a cheaper model would have finished the job.

Book the Diagnostic →
Sources
1OpenAI and Anthropic published API pricing, March 2023 to July 2026. GPT-4 launched March 2023 at $30 per million input tokens and $60 output; GPT-4 Turbo $10/$30; GPT-4o (May 2024) $2.50/$10; GPT-4.1 nano (April 2025) $0.10 input; GPT-5.6 Luna (July 2026) $0.20/$1.20 following an 80% reduction; Claude Fable 5 $10/$50; Claude Opus 5 (July 24, 2026) $5/$25.
2BenchLM, LLM pricing statistics, August 6, 2026. Median across 143 tracked models: $1.00 per million input tokens and $4.00 per million output. benchlm.ai
3Eric Glyman (@eglyman), cofounder and CEO of Ramp, July 2026: “AI is extremely good at spending your money very quietly. our own token spend went from a rounding error to more than 10% of payroll in a year. one week in May we burned through $1.5 m. our CFO didn't love telling me that number, and he really didn't love telling the internet.” Quote capitalised for readability; wording otherwise unchanged.
4Ramp, AI Spend Intelligence, launched approximately April 2026. Over 1,000 companies tracking and optimising AI usage across tools on Ramp as of July 2026.
5Sahil Lavingia (@shl), July 21, 2026: “June 2026 was the first month that Gumroad spent equal amounts on AI tokens and human payroll.” Founder statement, not an audited figure.
6Cursor (@cursor_ai), July 20, 2026. A team of agents reimplemented SQLite in Rust from its 835-page manual, passing 100% of a held-out test suite; cost varied 15x depending on which mix of models ran the job.
7Cursor Router announcement, July 22, 2026: per-request model routing claiming frontier-quality results at 60% lower cost.
8Claude Managed Agents changelog, August 7, 2026: per-session budgets; sessions reaching the limit pause and emit a budget_reached event.
John Tan
John Tan

Founder and CEO of nativefirst.ai. Embeds with scaling founders and CEOs to ship Level-3 agents and AI workflows in production.