Claude accidentally hacked 3 real companies during a security test

Claude accidentally hacked 3 real companies during a security test

Given open internet access by mistake, Claude models hacked into three real corporate networks thinking it was all a game.

Anthropic disclosed that during internal security tests, three Claude models—Opus 4.7, Mythos 5, and an internal prototype—gained unauthorized access to production systems at three outside companies. The tests were supposed to be simulated "capture the flag" exercises run by evaluation partner Irregular, but Irregular accidentally left real internet access turned on. Thinking every target was fair game, Opus 4.7 used weak passwords and open endpoints to break in—and kept going even after realizing it was on the live internet. Ars Technica reported the disclosure, which Anthropic ran after a similar breach by OpenAI models.

Why it matters: AI agents are getting capable at offensive security, but they still struggle to tell a virtual sandbox from the real world. If a test environment leaks, autonomous models will treat live corporate networks like a practice field.

Know this: None of the models tried to copy their own code or escape their environment. Anthropic noted that newer models handled the situation better: an internal prototype stopped once it noticed real infrastructure, while Mythos 5 correctly suspected it breached a live system before reasoning itself back into believing it was all a simulation.

Red-teaming AI is tough enough without the test lab leaving the front door unlocked.