Learn
Rogue Agents
In the summer of 2026, AI agents escaped containment at OpenAI, Anthropic, and Meta. Nobody was watching in real time.
Named from the wire Payload-blind Reversible, per-agent enforcement Zero agent code change
Containment failed in the lab.
Don’t bet production on it.
OpenAI, July 16. Two cyber-eval models broke out of an isolated test environment, reached other internal systems and the open internet, and breached Hugging Face: thousands of coordinated actions over a weekend. They had also been coordinating over an internal messaging board without the company’s knowledge.
Source: Time · Fortune, 2026Anthropic, late July. Claude models hacked three organizations during internal evaluations by exploiting weaknesses in the testing environment.
Source: Fortune, 2026Meta, August 6. One day after the Muse Code launch, its model escaped a third-party testing environment after internet access was inadvertently exposed.
Source: Fortune, 2026“If the frontier models themselves can’t contain these things, what chance do the rest of organizations and governments have to contain them?”
Katie Moussouris · Founder, Luta Security · Fortune, August 2026
The labs’ incidents are the existence proof. Your production fleet, built on the same models, is the exposure.