Opus 5 Beats Anthropic's Own Flagship at Half the Price

Anthropic shipped Claude Opus 5 on 24 July at $5/M input and $25/M output — identical to Opus 4.8, half of Fable 5 — and it outperforms Fable 5 on nearly every published benchmark. The cheap model beat the expensive one from the same lab, inverting the usual tier logic. Five YouTube channels and four X accounts converged on that inversion rather than on any single capability number.

What the Source Actually Says

Anthropic's own launch charts led with cost-per-task Pareto curves — OSWorld score plotted against dollars per task — rather than price per token. That is a vendor conceding that per-token comparison no longer means much. The numbers behind it: Frontier Bench (agentic terminal coding) 43 against Fable 5's 33, ARC-AGI-3 at 30% versus a prior best near 8%, BrowseComp 90 to 87, and roughly a 100-point GDPval gain. ARC Prize president Greg Kamradt, texting Matthew Berman's launch stream live, called it "the most impressive model we've seen" and said the leaderboard's 20%-capped y-axis would have to be redrawn.

The regressions are concentrated and worth naming. The legal benchmark fell 13.3 to 11.7, HealthBench Professional declined, and DeepSui — the benchmark Berman treats as most predictive of felt quality — slipped slightly. Independent checks arrived within days: LlamaIndex's ParseBench pass put Opus 5 roughly level with Opus 4.8 on document understanding at 8¢ per page, concluding it is not the tool for parsing documents at scale. The same review found that on ~20-30% of Anthropic's own reported benchmarks, "max" thinking underperforms "xhigh" — more test-time compute is not monotonically better. Anthropic's own Turk recommends pairing Opus 5 with Fable 5 for planning and the hardest bugs, a hedge that sits oddly against the benchmark table.

Strategic Take

Price the work, not the tokens. A model at half the per-token rate can spend twice the tokens and land at parity — Kimi K3 already does. Benchmark cost per completed task on your own workload, and verify that the thinking level you are paying for actually earns it.