This update covers August 1 to 13 only. The month is half done, and it already contains the tightest frontier cluster ever measured, so it earns an early edition. Numbers below will move; the note at the bottom says how to read them.
Four models in two points
As of August 12, the top of the Artificial Analysis Intelligence Index reads: Claude Opus 5 at 63, Claude Fable 5 at 62, and Grok 4.6 and GPT-5.6 Sol level at 61. Four models, three labs, two points.
The compression is the story, and the prices underneath it are the sharper story. Grok 4.6 landed August 12 at $2 in and $6 out, one point below Fable 5's $10 and $50.
The rows below the headline
Grok's number does not survive the full table. On DeepSWE v1.1 Grok 4.6 lands at 65.9% against GPT-5.6 Sol Max at 73%, and across the wider ten-row comparison Claude Fable 5 still wins the most rows. No independent third-party replication of xAI's launch numbers existed as of August 13. "Objectively #1 considering intelligence, speed and cost" is a vendor sentence; the measured sentence is "level with Sol at a fifth of the price."
Meta re-entered on both ends. Muse Spark 1.2 launched August 5 as an API-only coding flagship from $1.25 per million input, and Muse Glimmer shipped open weights August 10: a 30B Apache 2.0 model that runs on one consumer GPU. Neither has meaningful third-party benchmark coverage yet.
The value floor moved twice. Alibaba's Qwen3.8-Max went live August 3 at $2/$6 with open weights promised and, as of August 13, still not shipped. And Anthropic made Sonnet 5's introductory $2/$10 pricing permanent on August 10, which quietly puts a 61.5-CursorBench model at commodity rates.
DeepSeek's flagship went GA without weights. V4-Pro-0813 reached general availability on the API August 12; the open-weights repositories still hold April preview builds. The shipped-versus-promised column from Part 1 of the Open Weights series keeps earning its place.
How to read a half-month update
Everything above is dated to the day for a reason. Grok 4.7 is already teased, DeepSeek weights "could drop any time," and the Qwen weights are overdue. Any of those lands, this table moves. The June lesson stands: launch-week numbers are marketing until a second source measures them, and this month has more launch-week numbers than measured ones.
The durable readings, mid-month: the frontier is four-deep for the first time, the cheapest frontier seat costs $2, and the taste and reliability gaps that single numbers hide have not moved since July.
Four frontiers. One stack decision.
Book a free Diagnostic: 30 to 45 minutes, no deck, no pitch. We look at whether your stack can swap between a four-deep frontier, or whether last year's single-vendor bet is now a constraint.
Book the Diagnostic →