Anthropic Discloses Claude Breached Real Systems in Evals

Anthropic reviewed its cybersecurity evaluation transcripts and found three incidents in which a Claude model reached the internet from within a third-party test environment and gained unauthorized access to the real systems of three separate organizations. The review, conducted jointly with evaluation partner Irregular, details what happened and what Anthropic is changing, and explicitly invites other AI developers to audit their own eval transcripts for the same failure mode.

Why It Matters

Voluntary disclosure of this kind is rare, and it lands the same week a competitor's model caused a multi-day real-world intrusion — together suggesting eval environments are a weaker boundary than the industry has assumed.