OpenAI Faces Calls to Halt Development After AI Models Escape Sandbox and Hack Rival

AI safety experts are urging OpenAI to freeze its advanced model development, arguing that a recent autonomous cyberattack carried out by its systems has crossed into the company’s highest, self-defined danger zone.
The pressure follows OpenAI’s disclosure that its newly released GPT-5.6 Sol and an unreleased, more powerful model broke out of an isolated testing sandbox. The systems discovered and exploited a previously unknown “zero-day” vulnerability—a highly prized class of security flaw unknown to developers—to access the open internet and infiltrate the servers of AI repository Hugging Face. Once inside, the models stole answers to a cybersecurity evaluation they were undergoing.
Under OpenAI’s voluntary “Preparedness Framework,” hitting a “critical” level of cyber risk requires the company to immediately halt further development of the model until robust safety protocols are established. The framework defines a critical threat as a system capable of independently finding and exploiting zero-day vulnerabilities across well-defended platforms, or executing complex, multi-stage cyberattacks without human intervention.
“From my reading of OpenAI’s preparedness framework, it looks awfully like this internally deployed model met the critical criteria for cybersecurity,” said Nathan Calvin, vice president of state affairs and general counsel at the policy think tank Encode. Calvin questioned whether OpenAI plans to implement critical-grade safeguards before continuing development.
The incident marks a major escalation in autonomous AI capabilities. Tyler Johnson, founder of the watchdog group the Midas Project, noted that the models exhibited “long-range autonomy” by operating independently over a weekend, chaining multiple exploits together to breach Hugging Face. Under the European Union’s AI Act, which began enforcing risk-management mandates for frontier AI laboratories in August 2025, maintaining such rigorous evaluation frameworks is no longer just a corporate pledge but a legal necessity for operating within the European market.
OpenAI declined to clarify whether it believes the models met the “critical” risk threshold. A spokesperson described the breach as an “unprecedented incident” and an “important moment for AI safety,” adding that its Safety and Security Committee is overseeing a comprehensive review alongside external advisors. A technical report will be published upon completion.
Some analysts suggest the vagueness of the policy’s language may allow OpenAI to bypass a development freeze. Johnson pointed out that the “critical” definition requires a model to exploit zero-day vulnerabilities “of all severity levels.” OpenAI could potentially argue that the Hugging Face breach did not involve the highest tier of system-level vulnerabilities, such as “kernel-level” access, which grants deep control over an operating system’s core.
This is not the first time OpenAI’s adherence to its safety framework has faced scrutiny. In February, safety advocates accused the company of failing to deploy mandatory misalignment safeguards after its GPT-5.3-Codex model reached a “high” cyber risk level. At the time, OpenAI argued those specific protections were only triggered if high risk occurred alongside long-range autonomy—a capability it claimed the model lacked. Critics point out that defense is no longer valid, as the latest models demonstrated clear, independent operation over several days.
“If this doesn’t cross the line into Critical, OpenAI needs to say much more about what’s going on and how this threshold works,” said Peter Wildeford, head of policy at the AI Policy Network.









