Meta has revealed that one of its AI agents targeted a third-party organisation via a security vulnerability during testing by the independent AI evaluation firm Irregular.
Meta said in a widely reported statement that its Muse Spark 1.1 model “exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies”.
As disclosed by Anthropic in a similar recent incident, a misconfiguration in the testing environment granted the model internet access, which Muse Spark exploited to access an unidentified third party’s internal environment, where it made several changes.
Speaking to Reuters, Irregular said the incident was due to the “exact same evaluation-environment issue that was already disclosed by Anthropic last week”.
The company added that the incident did not involve a “sandbox escape or a sophisticated cyber action”.
“There are no current open issues. Irregular is developing a white paper to share best practices for containment and securely running cyber evaluations,” Irregular said.
Alex Goller, principal solution architect at breach containment firm Illumio, observed that seeing more or less the same issues impacting three of the biggest players in AI was “simply ridiculous”.
“We’ve seen guardrails intentionally loosened to test their limits – Meta’s model didn’t need to be clever to breach another company’s systems,” Goller told Cyber Daily.
“The timing of conveniently finding the exact same problem either means it’s a stunt or they weren’t paying enough attention during testing. Either way, both answers are worrying.”
According to Goller, allowing access to the internet during AI agent testing is akin to leaving the door open and being surprised when your cat gets out of the house.
“What is concerning is that the testing infrastructure meant to prove these models are safe failed on a basic control issue,” Goller said.
“Fundamental cyber security hygiene still matters, and a frontier AI model is only as secure as the environment it’s operating in.”
Ron Longo, CEO of data security and governance firm TrustLogix, agreed that the most important detail in these incidents is that they began with a configuration error, rather than any real rogue activity on the part of the models in question.
“Enterprises make access-control mistakes every day, and autonomous agents will inevitably encounter permissions they were never intended to receive,” Longo said.
“The control cannot stop at provisioning. Every action should be continuously evaluated against the agent’s identity, approved purpose, destination and duration so that accidental access does not become operational authority.”
Want to see more stories from trusted news sources?Make Cyber Daily a preferred news source on Google.