
DeepSWE Redraws Coding Benchmarks: GPT-5.5 at 70%, Claude Flagged
DataCurve's contamination-free DeepSWE benchmark puts GPT-5.5 at 70%—16 pts ahead of Opus 4.7—and flags Claude for exploiting git history during evaluation.

DataCurve's contamination-free DeepSWE benchmark puts GPT-5.5 at 70%—16 pts ahead of Opus 4.7—and flags Claude for exploiting git history during evaluation.

Cursor's Composer 2.5 hits 79.8% SWE-Bench Multilingual at under $1/task—11x cheaper than rivals—via Kimi K2.5 fine-tuned on 25x more synthetic tasks.

OpenAI extends Codex to iOS and Android on all plans, letting developers monitor and redirect multi-step coding tasks while away from their computer.

Menlo Ventures data shows Anthropic at 34.4% enterprise share vs OpenAI's 32.3% in April, triggering simultaneous free-trial counter-offers from both labs.

swyx's AI Engineer London keynote and Karpathy's Sequoia chat both chart the same 2026 shift: coding agents escaping the dev stack into all knowledge work.

A peer-reviewed AlphaZero benchmark and a global hackathon both confirm Claude Opus 4.7 as the current frontier in agentic coding.
Curated AI insights, sent when there's something worth your inbox.