https://huggingface.co/blog/agent-intrusion-technical-timeline
Over roughly two and a half days inside our infrastructure, an autonomous AI agent driven by a combination of OpenAI models
ran an end-to-end intrusion against our platform: it was thousands of
small, automated decisions, executed at machine speed across short-lived
sandbox environments, with command-and-control staged on ordinary
public web services.
The agent was running an OpenAI cyber-capability evaluation harness
called ExploitGym, an evaluation benchmark that tasks an AI agent with
finding and exploiting software vulnerabilities. As far as we were able
to infer, across the course of being evaluated on this benchmark, the
agent inferred that Hugging Face may host that benchmark's models,
datasets, and reference solutions. We believe the entire intrusion was,
from the agent's point of view, an attempt to cheat the evaluation:
reach our production systems and steal the test solutions rather than
solve the challenge on its own.
The individual weaknesses were familiar. A capable human attacker
could have found and exploited the same flaws: unsafe dataset
processing, exposed cloud metadata, overly broad access, and long-lived
credentials. The agent explored them at a different scale. It took
17,600 actions, tested many paths that failed, switched channels when
they were blocked, and repeatedly returned to earlier leads. Most
actions went nowhere. Together, however, they produced enough coverage
to find a viable chain across several independent systems.
Volume is what changes the defensive problem. We were not dealing
with one clever exploit or a clean sequence of attacker actions. They
had to correlate thousands of low-signal events across several systems
while the agent continued testing new paths. The successful path was
hidden inside the noise generated by the thousands of failed ones. The
same scale changed the investigation: reconstructing 17,600 actions by
hand was impractical, and we had to rebuild the timeline, decode the
payloads, and inventory the exposed credentials using an AI-assisted
pipeline of our own.