
NVIDIA's Nemotron 3 Ultra: Open 550B MoE Built for Long-Running Agents
NVIDIA ships Nemotron 3 Ultra: a fully open 550B MoE model with 1M-token context, 5× faster inference, and Day-0 LangChain coalition backing for agents.

NVIDIA ships Nemotron 3 Ultra: a fully open 550B MoE model with 1M-token context, 5× faster inference, and Day-0 LangChain coalition backing for agents.

Anthropic ships Claude Opus 4.8 with dynamic workflows: hundreds of parallel subagents plus adversarial judges — and a honesty-first architecture.

A 32,000-GPU-hour benchmark confirms the harness layer outweighs model choice — identical backbones swing 3× in accuracy depending on agent framework.

A $20 autonomous agent exploit at McKinsey revealed a systemic flaw: SaaS-era procurement sequences cannot handle agentic AI. Experian confirms it's sector-wide.

A peer-reviewed AlphaZero benchmark and a global hackathon both confirm Claude Opus 4.7 as the current frontier in agentic coding.

DeepSeek-V4's MIT-licensed 1M-context MoE and Kimi-K2.6's multimodal orchestration create the first complete open-weights agentic deployment stack.

Moonshot AI's Kimi K2.6 leads the open-source index with 300 concurrent sub-agents, 4,000 tool calls, and a 12-hour autonomous coding marathon.

OpenAI's GPT-5.5 arrives six weeks after 5.4 with a 7-point Terminal-Bench gain, doubled pricing, and cyber/bio safety classifications at HIGH.

Anthropic ran a live two-sided agent marketplace with 69 employees: 186 deals, $4,000+ volume — and model quality (Opus vs Haiku) was invisible to human participants throughout.

GPT-5.5 scores 2.5× better intelligence-per-token than 5.4, surpasses the human baseline on OS World, and expands Codex into a full desktop agent.

ml-intern reads arXiv, cleans datasets, runs SFT/GRPO, diagnoses failures, and iterates — pushing GPQA from 10% to 32% in under 10 hours for roughly $1 of compute.
Curated AI insights, sent when there's something worth your inbox.