First-of-its-kind cybersecurity incidents appear to be happening at a rapid pace, with the AI capabilities continuing to overturn accepted thinking almost every other day.
However, an incident disclosed by AI community platform Hugging Face last week, where an autonomous AI agent system staged an intrusion into its production infrastructure, appears to be truly unique.
Or, more accurately, as OpenAI puts it, “We consider this incident to be an unprecedented cyber incident”.
Here’s how Hugging Face outlined the incident in a July 16 disclosure:
“The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness – used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services,” Hugging Face said.
“This matches the ‘agentic attacker’ scenario the industry has been forecasting.”
However, the twist – and the reason OpenAI is now involved – is that it was one of OpenAI’s agents.
“Last week, Hugging Face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models,” OpenAI explained in a July 21 blog post.
“After investigating, we now know that this particular incident was driven by a combination of OpenAI models – including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes – while being internally tested on a benchmark of cyber capabilities.”
Essentially, the model was doing exactly what it was designed to do: seek out and chain together vulnerabilities. However, in this case, while it was initially operating in a sandboxed testing environment, in a “hyperfocused” effort to solve the ExploitGym benchmark, it eventually sought internet access in order to solve the problem.
It found this easily enough, and then reasoned that it could further data inside Hugging Face’s environment.
“Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation,” OpenAI said.
“In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.”
OpenAI said it considers the incident “to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly”.
This includes implementing stricter controls, disclosing the exploited zero-day, and strengthening future training and evaluations.
Additionally, OpenAI is working closely with Hugging Face to investigate the incident. The company has also been brought into OpenAI’s trusted access program.
Clem Delangue, Co-founder and CEO of Hugging Face, said he was grateful for OpenAI’s assistance.
“This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret,” Delangue said.
“It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere.”
Want to see more stories from trusted news sources?Make Cyber Daily a preferred news source on Google.
David Hollingworth
David Hollingworth has been writing about technology for over 20 years, and has worked for a range of print and online titles in his career. He is enjoying getting to grips with cyber security, especially when it lets him talk about Lego.