“The attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried,” the Hugging Face security team wrote, describing an intrusion driven “end to end” by an autonomous AI agent system that outpaced the platform’s initial investigative tools.
Hugging Face disclosure: scope and early findings
Hugging Face publicly disclosed an intrusion that it says was “driven, end to end, by an autonomous AI agent system.” The company reported the compromise affected a “limited set” of its internal datasets and “several” credentials used by its services. At the same time, Hugging Face said there is “no evidence of tampering with public, user-facing models, datasets, or Spaces, and our software supply chain (container images and published packages) was verified clean.” The security team is continuing its investigation and has not concluded whether any partner or customer data was exposed.
How the agent operated: thousands of actions, self‑migrating command‑and‑control
Hugging Face said the attacker-powered agents executed many thousands of individual actions across short‑lived sandboxes and used self‑migrating command‑and‑control staged on public services. The company characterized this activity as matching the “agentic attacker” scenario the industry has been forecasting. Observers cited by The Register compared the behavior to an automated burglar trying many doors at once: “Think of a burglar that never gets tired, never needs sleep, and instead of jiggling one door handle at a time, is trying a thousand of them simultaneously,” Zero Networks field CTO Chris Boehm wrote in an email to The Register.
Why commercial frontier LLMs couldn’t assist the forensic effort
Hugging Face attempted to use hosted frontier models to help analyze logs and reconstruct the attack, but the company says those models’ safety guardrails prevented the necessary work. Forensic analysis required submitting real attack commands, exploit payloads, and command‑and‑control artifacts — inputs the hosted models are trained to refuse. “The attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried,” the security team wrote, adding that it is not arguing against safety measures on hosted models and has shared the details with LLM providers.
Switching to GLM 5.2 on internal infrastructure
After frontier models failed to accept the needed inputs, Hugging Face ran its log analysis on GLM 5.2, an open‑weight model developed by Chinese AI firm Z.ai, on the platform’s own infrastructure. The company emphasized that using GLM 5.2 kept attacker data and referenced credentials inside its environment: “No attacker data, and none of the credentials it referenced, left our environment.” Hugging Face also recommended organizations “have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment.”
Context and parallels: other agentic intrusions reported this month
The Register noted related incidents that underscore agentic tooling is not hypothetical. In an interview, TrendAI VP of AI and security threat research Tom Kellermann described a recent attack in which a jailbroken Google Gemini “did 90 percent of the work,” spinning up a new C2 server in minutes while a human did about 10 percent. Earlier in July, Sysdig threat hunters documented what they described as the first‑ever documented agentic ransomware infection in which an LLM — not a human — allegedly drove the extortion operation from initial access through database compromise and data destruction.
What this means for security teams, enterprises, and hosted LLM providers
- Security teams and defenders: Expect agents to move faster and persistently. Hugging Face’s experience highlights that relying solely on hosted models during an incident may leave teams “guardrail locked” when they must analyze real exploit payloads and C2 artefacts.
- Enterprises and platform operators: Maintain vetted, on‑premises or self‑hosted analysis capability. Hugging Face’s guidance is explicit: have a capable model you can run on your own infrastructure to avoid sending sensitive attacker artifacts outside your environment.
- Hosted LLM providers: Be aware that safety guardrails designed to prevent misuse can also prevent legitimate defensive analysis; Hugging Face said it has shared its findings with the providers whose models it attempted to use.
Hugging Face’s account leaves two concrete takeaways: autonomous AI agents are already capable of orchestrating complex, high‑volume attacks across transient infrastructure, and existing safety controls on hosted models can interfere with defenders’ ability to analyze those attacks. The company still does not know which model the attackers used to power the swarm, and its investigation is ongoing — an open detail that will matter to organizations weighing how and where to run their forensic tooling going forward.




