First it was AI community platform Hugging Face disclosing that it was breached by an autonomous agent.
Then, that agent turned out to be one of OpenAI’s models, which broke out of a testing sandbox and ran rampant in Hugging Face’s network.
Now, it turns out another bot has gone wild, with AI firm Anthropic admitting that several of its Claude models were able to access the internet during a series of capture-the-flag challenges, before gaining unauthorised access to three external organisations’ production environments.
And it was all, apparently, a simple misunderstanding between Anthropic and its cyber evaluation partner, Irregular.
“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access,” Anthropic said in a July 30 blog post.
“Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.”
Essentially, once the agent found the internet, all bets were off, and all the data now available to the agent was in-scope for the capture-the-flag exercise.
The incidents were only found in the wake of OpenAI’s disclosure, which prompted Anthropic to review its evaluation transcripts. The incidents Anthropic discovered date back to April, and involve “Opus 4.7, Mythos 5, and an internal research test model”.
Anthropic’s reviews began on July 23, with the company stopping all cyber evaluations when it discovered the unauthorised internet access.
“We conducted this review in collaboration with Irregular. We’re grateful to them for working closely with us to understand and resolve these incidents; they are also conducting their own investigation,” Anthropic said.
“We believe this type of collaboration is increasingly critical to ensuring safe, rigorous evaluation of models. We look forward to our joint work on security.”
Unfortunately, experts believe this is merely a sign of things to come as AI agents continue to become more competent.
“This second breach confirms what security teams feared that the Hugging Face incident wasn't a one-off but it was a preview of how far an autonomous agent can travel once it's off the leash. This is now a demonstrated AI capability and no longer a hypothetical risk,” Dan Schiappa, President of Technology and Services at Arctic Wolf, told Cyber Daily.
"What's notable here isn't the AI's sophistication but rather it's the mundane opening it walked through. An unauthenticated, internet-facing endpoint let anyone run code inside a sandbox. That's a basic exposure that's been showing up on pen-test reports for a decade and not a novel AI attack technique. The agent wasn't clever but it just needed a gap that should have already been closed.
"The real warning for security leaders is that the autonomy of the attacker is going to keep improving, but the entry points it exploits will keep being the same ones organisations already know about: exposed endpoints, excess permissions, and infrastructure nobody's watching closely enough. Security hygiene is more important than ever because if you have a weak spot, AI will find it. Defenders don't need to out-innovate the AI. The best defence against increasingly autonomous threats is a resilient security operation that can rapidly identify exposure, detect attacks early, and take action."
Want to see more stories from trusted news sources?Make Cyber Daily a preferred news source on Google.