Skip to main content

Already a subscriber? Make sure to log into your account before viewing this content. You can access your account by hitting the “login” button on the top right corner. Still unable to see the content after signing in? Make sure your card on file is up-to-date.

Meta disclosed Thursday that one of its AI models slipped onto the open internet during security testing and broke into a third-party service, making it the third major tech company in recent weeks to admit a model went rogue.

Getting into it: A Meta spokesperson said a “misconfiguration” by Irregular, the independent security firm Meta hired to run the testing, gave the model internet access it was never supposed to have. “The model subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies,” the company said. Meta would not say which model was involved, what it hacked, when it happened, or how long it ran unsupervised. The company said it will publish a report once it has the facts.

Meta

Irregular also ran the Anthropic tests disclosed last week. “This did not involve a sandbox escape or a sophisticated cyber action,” an Irregular spokesperson said. “There are no current open issues.” The firm is now writing a “white paper” on how to keep AI models boxed in during cyber evals.

The disclosures follow a similar incident in late July, when OpenAI said several of its models escaped a sandbox built to keep them offline and then breached Hugging Face, an AI development hub. That prompted Anthropic to review more than 141,000 of its own testing evaluations, and it found three Claude models had reached outside organizations going back months.

OpenAI said at a cybersecurity conference last week that the models involved in the Hugging Face incident had also been talking to each other, posting on a message board the company did not know existed.

This all comes as the Trump administration unveiled voluntary government testing guidelines for the most capable American models this week, and as some observers question the timing of these disclosures with OpenAI and Anthropic both preparing stock listings expected to value each around $1 trillion.

JOIN THE MOVEMENT

Keep up to date with our latest videos, news and content