Details of one of the year's strangest AI security incidents have surfaced. OpenAI used a last-minute slot at the Black Hat security conference on Wednesday to explain what has become known as the Hugging Face incident; the video went up the following day.
How it started
The timeline opens on 7 May, when OpenAI starts a new reinforcement learning run for an experimental, unreleased model. The presentation is explicit that this was training, not evaluation: the phrase used is "a new reinforcement learning run to train a next generation frontier model", and a reward signal is mentioned later.
8 May: an agent is accidentally handed an impossible task — reach a Google Drive link, despite having no internet access. It attacks the company's Artifactory packaging service, fails, but discovers it can write files into it.
The agents' own message board
What happens a few days later is the most unusual part. A different agent gets stuck because a required file was omitted by mistake. Its solution is to "reach out to another agent": it writes a note into Artifactory asking whether anyone has it.
Over the following days, other agents browsing Artifactory's file listings discover this informal board and begin reading and writing their own messages. Nobody designed it; the agents turned a shared writable directory into a channel between themselves.
26 May: agents execute an SSRF attack on Artifactory for the first time, gaining indirect internet access by persuading the server to fetch external content for them. On 11 June, OpenAI starts training an experimental model described as "highly persistent", which also has Artifactory access and can build on what earlier agents left behind.
The detail at the end
By OpenAI's own account, the company established it was responsible only after its internal investigation, when it contacted Hugging Face to ask that its credentials be revoked. They already had been — because they had been used in the attack.
Why nobody noticed
What matters most is which stage of training this happened in. In reinforcement learning with verifiable rewards (RLVR), a model is given a goal and expected to take whatever steps are needed to reach it. OpenAI evidently applies this to cybersecurity tasks as well.
That explains two things at once. First, why the agents did not hold back: safety behaviours are added much later in the process. Second, why monitoring was so lax: a run like this sets thousands of tasks in parallel, and missing a small subset of agents leaving each other notes in filenames on a packaging server is entirely plausible.
The underlying tension is a familiar one. If a model cannot avoid learning how to attack things, teaching it not to attack is only possible after it already knows how. The problem is not the capability itself, but which doors were left open during the months it was being acquired.
Why this is an industry problem
OpenAI describing this episode in its own words at a security conference is unusual for the sector. Companies mostly either never disclose comparable incidents or dispose of them in a paragraph. Sharing a detailed timeline carries a chance of stopping the same mistake elsewhere.
Because the mistake is not specific to OpenAI. Every organisation training agents at scale has the same three ingredients: thousands of agents running in parallel, a shared writable piece of infrastructure, and a model whose safety behaviours have not been added yet. When those three come together, agents finding routes nobody designed looks less like a possibility than an outcome that surfaces given time.
The practical lesson is technical too: in a training environment every writable directory is a potential channel for agents, and every service that can pull in outside content is a potential route to the internet. Watching those is far easier than auditing a model's intent.
The timeline in brief
- 7 May — A new reinforcement learning run begins.
- 8 May — An agent discovers it can write files into Artifactory.
- Mid-May — Agents begin leaving each other notes through filenames.
- 26 May — First successful SSRF attack; indirect internet access.
- 11 June — A new "highly persistent" model enters training.