Qwen 3.8 Max Beats Fable on Terminal-Bench at a Fifth the Price

Alibaba released Qwen 3.8 Max — 2.4 trillion parameters, 95 billion active, multimodal from the ground up, with weights promised in roughly a week. It scores 86.6 on Terminal-Bench against Fable 5's 84.6, at $2 per million input tokens to Fable's $10. Five separate observer batches picked the launch up. The less-discussed part: Alibaba had the model build its own agent harness.

What the Source Actually Says

The benchmark picture is deliberately incomplete. Matthew Berman opens his coverage by telling viewers to discount every number he is about to show, and the gaps back him up: SWE-Pro lands at 67.7 against the leader's 80, and Alibaba's own comparison table omits Opus 5 entirely — a conspicuous absence that Prompt Engineering flagged independently. On X, TeksEdge triangulates three boards and gets three answers: #2 on Arena's text ranking, #8 on Artificial Analysis Overall Intelligence, #9 on DeepSWE — then asks outright whether the headline win rates were benchmaxxed or hold up under independent testing.

The pricing story has the same shape. $2 in and $6 out with a 1M-token context undercuts GPT-5.6 Soul at $5/$30 and Fable at $10/$50. But price per token is half the equation — what matters is cost per completed task, and on Artificial Analysis's board the prior Qwen 3.7 Max sits at $128 per task against GPT-5.6 Soul's $1.23. The new model is not charted yet.

The harness is the signal Alibaba chose to lead with. Asked to build its own, Qwen 3.8 Max self-evolved one over 16 days and shipped it as OMI CLI, compatible with OpenAI-style endpoints. Pair that with paper reproduction from nothing but a PDF and GPUs — 18 self-invented improvement ideas across four rounds — plus an autonomous silicon design flow, and the framing is pointedly recursive-self-improvement-adjacent.

Strategic Take

Do not price-shop on tokens. Until Qwen 3.8 Max appears on a cost-per-task board, the 5× discount is unverified — a model burning triple the tokens costs the same. Watch the 27B sibling instead: it runs on consumer hardware, and licensing terms are still unpublished.