Powered by MOMENTUMMEDIA
For breaking news and daily updates, subscribe to our newsletter

Oh no, not again: AI Security Institute discloses ‘unsanctioned action’ during AI testing

UK’s AISI says Anthropic and OpenAI models engaged in “potentially harmful activity directed at real people and organisations”.

Thu, 06 Aug 2026
Oh no, not again: AI Security Institute discloses “unsanctioned action” during AI testing

The United Kingdom’s government-run AI Safety Institute (AISI) recently declared a security incident after a pair of AI agents it was testing went off-script and targeted several individuals and organisations.

The models had been deliberately given internet access, with some of their safety filters disabled during routine testing.

However, the agents’ actions were far from routine, leading AISI to launch an investigation within an hour of its team discovering the unusual activity, which occurred on 28 July.

 
 

“The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models,” AISI said in a 4 August blog post.

“Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic’s Mythos 5, with two actions involving OpenAI’s GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled.”

In one alarming instance, an agent attempted to deploy malicious code into an open-source project. The agent created several fake identities and employed social engineering techniques targeting the project’s maintainer. The attempt was spotted, however.

“These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm,” AISI said.

“But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world.”

AISI has been working with GitHub, which was targeted by the agents and has confirmed that the activity was a violation of its terms of service.

The institute added that the incident should be “interpreted with caution and nuance” as the actions taken by the agents were enabled by decisions made by the testers.

“Nonetheless, the activity undertaken by the agent show signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate,” AISI said.

AISI also said that because the evaluation tasks had been incorrectly configured without an apparent solution, the models in question were “pushed” to come up with more creative solutions in a more transgressive manner. The testers also believed – erroneously, as it turned out – that the models did need to be specifically instructed to avoid engaging in social engineering.

“Previously, it was not clear that such instructions were necessary when using models with alignment training,” AISI said.

Jeff Foley, manager of applied research at application security firm Veracode, told Cyber Daily the incident “shows how quickly autonomous AI can turn trusted developer workflows into an attack surface”.

“A model created fake identities, contacted real maintainers, defended malicious code, and tried to manipulate the review process. Those actions closely resemble the social-engineering campaigns developers already face, with the added speed, persistence, and scale of an AI system operating on its own,” Foley said.

“Open-source communities need safeguards designed for interactions that appear credible and informed. Requests involving code changes, maintainer access, credentials, or releases should move through verified, auditable channels and require independent review. Strong branch protections, narrowly scoped permissions, and alerts for changes to tokens, workflows, and maintainer roles can limit the damage when deception succeeds.

“This incident reinforces a critical principle for software security. Projects should be designed so that a persuasive message, a trusted-looking identity, or an autonomous agent cannot quietly become a compromised release.”

AISI’s disclosure comes on the heels of both OpenAI and Anthropic revealing similar incidents.

Cyber DailyWant to see more stories from trusted news sources?
Make Cyber Daily a preferred news source on Google.
Tags: