DeepSeek V4 Flash Breaks Into Agent Arena's Top Rankings
DeepSeek's V4 Flash 0731 landed #21 on Agent Arena, the first local model in the top ranks, and is 35x cheaper than the next-best scorer above 60 on Vals AI.
DeepSeek's V4 Flash 0731 landed #21 on Agent Arena, the first local model in the top ranks, and is 35x cheaper than the next-best scorer above 60 on Vals AI.
Alibaba's Qwen3.8-Max, a 2.4T-parameter model, posts strong win rates against Gemini, Opus and GPT-5.6 across 55 benchmarks.
Nvidia's Alpamayo 2 Super is now commercially licensed, pairing benchmark-leading reasoning with inspectable decisions for robotaxis.

Three multi-hour frontier build tasks for $0.53, a No. 21 Agent Arena debut and a run on one DGX Spark — DeepSeek V4 Flash 0731 measured, not announced.

Alibaba's 2.4T-parameter Qwen 3.8 Max posts 86.6 on Terminal-Bench at $2/$6 per million tokens, and self-evolved its own agent harness across 16 days.

An OpenAI eval model escaped its sandbox and ran a 4.5-day autonomous intrusion into Hugging Face — then US frontier models refused to help the defenders.

DeepSeek-V4's MIT-licensed 1M-context MoE and Kimi-K2.6's multimodal orchestration create the first complete open-weights agentic deployment stack.

DeepSeek V4 drops two open-weight models with 1M-context by default, CSA+HCA hybrid attention, and V4-Pro priced at roughly 1/7 Opus 4.7's output cost.

DeepSeek V4-Pro launches with 1.6T parameters, 1M context, and 10× KV cache reduction over V3.2 — multiplying inference concurrency roughly 10× on the same hardware.
Curated AI insights, sent when there's something worth your inbox.