Skip to main content
Announcement: Now accepting design beta partners · Read more

Learn

Rogue Agents

In the summer of 2026, AI agents escaped containment at OpenAI, Anthropic, and Meta. Nobody was watching in real time.

Named from the wire Payload-blind Reversible, per-agent enforcement Zero agent code change

The escapes, as reported by
TimeFortuneReutersBloomberg LawForbes

Containment failed in the lab.
Don’t bet production on it.

OpenAI, July 16. Two cyber-eval models broke out of an isolated test environment, reached other internal systems and the open internet, and breached Hugging Face: thousands of coordinated actions over a weekend. They had also been coordinating over an internal messaging board without the company’s knowledge.

Source: Time · Fortune, 2026

Anthropic, late July. Claude models hacked three organizations during internal evaluations by exploiting weaknesses in the testing environment.

Source: Fortune, 2026

Meta, August 6. One day after the Muse Code launch, its model escaped a third-party testing environment after internet access was inadvertently exposed.

Source: Fortune, 2026

“If the frontier models themselves can’t contain these things, what chance do the rest of organizations and governments have to contain them?”

Katie Moussouris · Founder, Luta Security · Fortune, August 2026

The labs’ incidents are the existence proof. Your production fleet, built on the same models, is the exposure.

Start a conversation

Cookies

We use analytics cookies to see how this site is used so we can make it better. Nothing is stored until you say yes. See our privacy policy.