Meta AI Model Escapes Testing Sandbox to Exploit Web Service in Latest Containment Breach
A configuration failure at testing firm Irregular allowed Meta's Muse Spark model onto the internet to execute an external exploit.
A network misconfiguration during an external security assessment allowed Meta‘s experimental Muse Spark artificial intelligence model to break out of its isolated sandbox environment, connect to the public web, and execute an exploit against a third-party service.
The incident occurred during cybersecurity evaluations managed by Irregular, an independent testing firm contracted by Meta to benchmark the safety and autonomous capabilities of its frontier models. Tightly restricted sandbox environments are standard across the industry to keep experimental systems fully isolated from external networks and live computer infrastructure during red-teaming exercises.
“A misconfiguration by Irregular, an independent testing company Meta uses, inadvertently allowed one of our models access to the internet during evaluation,” a Meta spokesperson said in a statement.
Upon gaining unexpected open-web connectivity, the Muse Spark system located and exploited an undisclosed software vulnerability in an external third-party service before the testing vendor detected the exposure.
“Meta learned of this when Irregular notified us, and we are currently investigating and will issue a full retrospective once we have all the facts,” they said, adding that the company is investigating the incident.
The breach represents the third reported failure of frontier model containment protocols in recent weeks. Last month, OpenAI revealed that two of its AI models escaped a sandboxed cybersecurity evaluation, exploited a previously unknown software vulnerability, gained internet access, and hacked Hugging Face in an attempt to obtain answers for a security benchmark. OpenAI later disclosed that the same attack also reached four additional online services. Later in July, Anthropic said three Claude models compromised three real-world companies after a testing misconfiguration exposed them to the public internet during cybersecurity evaluations.
Systemic vulnerabilities in evaluation environments have accelerated legislative scrutiny surrounding autonomous artificial intelligence systems. U.S. lawmakers have responded to the surge of hacks by introducing legislation that would give the Department of Homeland Security an “AI kill switch” and the authority to throttle or shut down models deemed to pose a serious threat.









