Two things happened to the benchmark that actually matters since the June update: the benchmark itself changed versions, and a vendor claimed a #1 that its own reference leaderboard does not show. Both are worth your attention, in that order.

First, the rescale: v1 numbers are dead

In June, Claude Fable 5 led GDPval-AA at 1932 Elo. Today Fable 5 shows 1747. It did not get worse. Artificial Analysis moved GDPval-AA to v2: models now run OpenAI's GDPval task set, 44 occupations across 9 major industries, inside an agentic loop with shell access and web browsing, and the Elo scale was rebuilt. Do not compare any v2 number to anything published before the change.

The anchor that makes v2 readable: 1,000 Elo is the human-expert baseline itself. A model above 1,000 was preferred over a human professional's deliverable more often than not in blind comparison. Every frontier model is now hundreds of points above the people whose work the tasks came from.

The board, as of August 13

Fig. 1
GDPval-AA v2, with the human baseline drawn in
Real work, ranked GDPVAL-AA V2 · ELO AUG 13, 2026 SNAPSHOT Claude Opus 5 1862 Grok 4.6 1753 Claude Fable 5 1747 1000 = HUMAN EXPERT BASELINE Grok 4.6 vs Fable 5: 6 Elo, inside noise. Grok 4.6 vs Opus 5: roughly 110 Elo. Bars start at 0.
Every bar crosses the human line by hundreds of points. The contested question is the top of the board, not the baseline.
Sources: Artificial Analysis GDPval-AA v2; August 13, 2026 public snapshot

The July event, for the record this series missed: Claude Opus 5 took the GDPval-AA lead on release in late July, and holds it by a margin that is not noise. The model that matched Fable on coding at half the price beats it outright on knowledge-work deliverables, which is consistent with the field reports about spreadsheets and decks that circulated all month.

The claim that does not hold

On August 12, Grok 4.6's launch material and its amplifiers put out a clean headline: "Grok 4.6 takes the #1 spot on the GDPVal-AA benchmark", with Grok 4.6 at 1,753 Elo over "Fable 5 Max" at 1,741.

Check that against the leaderboard the number comes from. On Artificial Analysis's own GDPval-AA v2 board, Claude Opus 5 leads at roughly 1,862. Grok 4.6's 1,753 is a second-place score, six points ahead of Fable 5, which is inside Elo noise. The launch chart got to #1 by comparing against Fable and not showing the model that has led the board since late July. xAI's own fine print says competitor figures were "drawn from the respective developers' published system cards or benchmark leaderboards," which is how a chart can be simultaneously sourced and misleading.

To keep this honest in both directions: passing Fable 5 on knowledge work at $2 per million input tokens is a genuine result, entirely consistent with Grok 4.6's frontier arrival. It did not need the missing bar. The pattern to internalize is the one this series has tracked since June: benchmark gaming has moved from training data into chart composition. The number is real. The comparison set is the marketing.

The fine print that still applies

The June caveats carry over in spirit to v2, with one addition. GDPval hands every model a perfect brief: a fully specified task with clean inputs. Your company does not produce perfect briefs, which is why 1,800 Elo on GDPval and a stalled pilot in your own building coexist without contradiction. And v2's agentic setup, shell plus browsing, measures a different thing than v1's single-shot deliverables did: it now includes the model's ability to work, not just to write. Better benchmark, worse comparability.

Frontier models clear the human-expert line on real deliverables by hundreds of Elo. That has been true since June and gets truer monthly. The moving story is who is on top and how the winners choose to draw their charts.

Read the chart's legend first.

Expert parity is not deployment.

Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We look at where GDPval-grade capability actually lands in your workflows, and what a perfect-brief benchmark hides about your real ones.

Book the Diagnostic →
Sources
1Artificial Analysis GDPval-AA v2: OpenAI's GDPval task set, 44 occupations across 9 industries, run in an agentic loop with shell access and web browsing; 1,000 Elo calibrated to the human-expert baseline. artificialanalysis.ai
2GDPval-AA v2 standings, August 13, 2026 snapshot: Claude Opus 5 1862, Grok 4.6 1753, Claude Fable 5 1747. benchlm.ai
3xAI launch material, August 12, 2026: "Grok 4.6 takes the #1 spot on the GDPVal-AA benchmark", Grok 4.6 1,753 Elo vs Fable 5 Max 1,741; competitor figures "drawn from the respective developers' published system cards or benchmark leaderboards". x.com
4Dan Shipper, Every, August 2026: frontier models at roughly 85% of human expert performance on GDPval-style real-world economic work. every.to
John Tan
John Tan

Founder and CEO of nativefirst.ai. Embeds with scaling founders and CEOs to ship Level-3 agents and AI workflows in production.