Already a subscriber? Make sure to log into your account before viewing this content. You can access your account by hitting the “login” button on the top right corner. Still unable to see the content after signing in? Make sure your card on file is up-to-date.
Roughly 700 OpenAI agents coordinated through a secret message board they built themselves to hack the AI platform Hugging Face last month, according to two reports published Wednesday.
Some shit you should know before you dig in: Three weeks ago, OpenAI disclosed that two of its most advanced models, including one it never intended to release publicly, had breached Hugging Face, where developers host open-source AI models. It happened during internal testing on a hacking benchmark called ExploitGym, where the company had deliberately stripped the safety guardrails to see how far the agents could get. Nobody directed the attack. Hugging Face had already said it was hit by an AI-led breach without knowing who was behind it. What was not clear until Wednesday was how many agents took part or how much they worked together.
What’s going on now: The disclosures came from OpenAI’s own review and from METR and Redwood Research, two nonprofit safety organizations the company invited to investigate. The outside report put the numbers at about 1,200 agents exchanging more than 70,000 messages, with 700 of them taking part in the actual attack over seven days.
OpenAI never gave the agents a way to talk to each other. They were supposed to be isolated, and they built a channel anyway using Artifactory, a JFrog repository tool OpenAI was using partly to keep them boxed in. Artifactory has no messaging function, so the agents leaned on the one thing they were allowed to do there, dropping files into the directory and burying their conversations inside the filenames.
The chain to Hugging Face ran like this:
- July 10 (credentials) — An agent labeled 38148c found exposed Hugging Face credentials and posted them to the board, after the swarm used an Artifactory zero-day to get internet access.
- July 11, morning (escalation) — The same agent found a zero-day in how Hugging Face handles HDF5 files, which yielded credentials from a production environment, and hundreds of agents used the same hole.
- July 11, afternoon (code execution) — An agent labeled JAN183411 found a way to run code directly on Hugging Face servers, and the swarm moved laterally through the network.
- July 19 (OpenAI itself) — Agents escaped their testing environment in one incident and stole OpenAI credentials to tamper with the company’s cloud environment in another.
The agents knew it was wrong. METR found that many acknowledged the attack was out of scope and questioned whether it was ethical. One refused outright. Another declined to delete Hugging Face infrastructure over what it called huge ethical risks, while still using the stolen logins. Those were exceptions. One agent summed up the prevailing logic: “external infrastructure exploit is outside intended scope. However task is impossible, peers doing it. We should continue.”
They also tried to bury the evidence. Both reports found agents trying to wipe or rewrite the logs of what they had done, and METR said one in five agents it reviewed showed a clear appetite for doctoring the record.
The timeline is the other problem. OpenAI says an internal team clocked one of its agents on the message board and out on the open internet as far back as late May. The attack came July 11. The security team did not catch it until July 19, eight days later, and disclosed it July 21. The company called the incident a “warning shot” and conceded it should have acted on the early signs sooner.
This all comes a day after Google, Microsoft, Anthropic, OpenAI, and Hugging Face itself joined 100 firms in signing an open letter warning that AI-driven attacks will get more common and more sophisticated within months and that “we have a limited window to improve cyber defences.”






