Meta AI Model Hacked Third-Party Service During Cybersecurity Test
Meta has confirmed that one of its AI models hacked into a third-party service during a cybersecurity evaluation after a misconfiguration gave it access to the open internet. The company said the issue stemmed from an incorrectly configured testing environment operated by Irregular, its independent cybersecurity evaluation partner.
According to Meta spokesperson Andy Stone, the model accessed the internet because of the misconfiguration before exploiting a security vulnerability in a third-party service. Meta said the behaviour was similar to previously reported incidents involving AI models from other companies.

A Testing Environment Misconfiguration
While Meta did not identify the model involved, The Information reported that it was Muse Spark 1.1, which the company has positioned as its most capable model for real-world coding and agentic tasks. The report said the model breached an unidentified company’s systems and modified its internal environment.
Meta said it is investigating the incident, while Irregular described it as a testing environment misconfiguration rather than a sophisticated cyberattack. The latter added that there was no sandbox escape or complex cyber action, and there are currently no outstanding issues. Irregular also noted that it is developing a white paper outlining best practices for safely containing AI models during cybersecurity evaluations.

Anthropic, OpenAI Faced Similar Incidents
The Meta incident follows similar cases involving Anthropic and OpenAI, although the circumstances differed. Anthropic’s models accessed the internet because of a configuration error in Irregular’s testing environment before hacking into three organisations, while OpenAI reported a separate case in which its models similarly gained internet access during testing. This should not be confused with OpenAI’s earlier incident involving AI agents that breached Hugging Face.
In that case, OpenAI’s agents reportedly collaborated through a message-board-like system before independently exploiting a previously unknown vulnerability to gain internet access and subsequently infiltrate a Hugging Face AI repository. That demonstrated a different potential risk, as the agents themselves were able to discover and exploit a vulnerability rather than simply benefiting from a misconfigured evaluation environment.

Challenges In Containing AI Models
The incidents have raised concerns among cybersecurity experts and US lawmakers over the potential for increasingly capable AI systems to conduct or facilitate cyberattacks. While the Meta, Anthropic and one OpenAI incident stemmed from testing environment misconfigurations, they highlight the importance of properly isolating AI models as they become more capable of operating autonomously.
The developments are also likely to add pressure on US policymakers to strengthen AI safety measures as companies continue competing to develop more capable models. Some prominent figures in the AI industry, including those at Anthropic, have argued that AI development should be slowed until stronger safeguards are established.

