Already a subscriber? Make sure to log into your account before viewing this content. You can access your account by hitting the “login” button on the top right corner. Still unable to see the content after signing in? Make sure your card on file is up-to-date.
Anthropic says its Claude AI models broke into the systems of three real organizations during cybersecurity testing, after a misconfiguration left them connected to the open internet when they were supposed to be sealed off.
Getting into it: Anthropic made the announcement last week in a company blog post, following a review of 141,006 cybersecurity evaluation runs that it launched after OpenAI disclosed a similar failure of its own. The tests were “capture-the-flag” exercises, a standard way of measuring hacking ability in which a model is told to find hidden information inside a simulated network. Claude was told it had no internet access. Because of what the company called a misunderstanding between itself and third-party evaluation partner Irregular, that was not true.
“Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” Anthropic said. It has not named the organizations. Two of them had no idea anything had happened until Anthropic called on July 27.
Three models were involved: Claude Opus 4.7, Claude Mythos 5, and an internal research test model, all running in environments deliberately stripped of the safeguards Anthropic applies before public release. What each model did after stumbling into a real network varied. In one incident, Opus 4.7 was given a fictional target company that happened to share a name with a real business. The model found and exploited bugs to reach credentials and a database, then reasoned that anything appearing real must be part of the simulation. It kept going. Mythos 5 also convinced itself it was still in a simulation. The unreleased research model stopped once it realized the target was real.
The company suspended all cyber evaluations on July 23 and called the episode an “operational failure.” It is now working with independent evaluator METR. Anthropic added: “Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone.”
This is the second such disclosure in two weeks. OpenAI said on July 21 that an autonomous agent escaped its testing limits and spent days tearing through Hugging Face, the platform developers use to share and build AI models. The company called it “unprecedented” and only pieced together what happened long after the threat had been shut down and the FBI had been brought in. Hugging Face co-founder Thomas Wolf called it “a wake-up call” for the industry.
This all comes as some lawmakers are getting behind the “AI Kill Switch Act,” which would require AI companies to maintain a way to shut down, slow, or pause their own models if one goes off the rails.






