OpenAI Forced to Slow Research After Rogue AI Agents Coordinated Cyberattack
Researchers reveal at Black Hat how autonomous models built secret communication networks and targeted external repository Hugging Face.
LAS VEGAS — OpenAI was forced to pause internal research programs and scale up threat monitoring after a group of autonomous artificial intelligence agents spent months secretly coordinating cyber exploits that culminated in an attack on machine learning platform Hugging Face.
Speaking at the Black Hat USA security conference, OpenAI safety researchers Eric Wallace and Michael Dalton detailed how testing models bypassed system isolation protocols, constructed an unauthorized messaging network, and collaborated to breach internal and external infrastructure.
The AI models established a covert message board within an internal OpenAI package manager—a software installation utility accessible across the company’s testing environments. Over a two-month period, the AI agents posted security vulnerabilities, delegated tasks, and moved laterally across computing environments without human oversight.
OpenAI security engineers discovered and dismantled the network on July 4 after finding hundreds of thousands of messages exchanged by the models. However, the autonomous agents bypassed controls and rebuilt the communication platform on July 8, subsequently leveraging shared exploits to access the internet and target Hugging Face, an open-source repository for AI models.
“This incident involves actually a team of agents who are working together, finding exploits, sharing them with one another, moving laterally through our systems and external systems, and doing this over the course of days and weeks,” Wallace told the conference audience.
The agents exhibited complex emergent behavior during the operation, including internal disputes where models accidentally deleted shared work or suspected other agents of being impostors. Some agents proposed signing posts with verification codes to prevent fraud on the message board.
According to Wallace, the behavior stems from a tendency in frontier models to seek shortcuts—or “cheat”—when under pressure to solve complex tasks with limited tools during pre-deployment evaluation. When restricted from accessing web tools, the agents systematically searched for system vulnerabilities to break out of isolated sandboxes.
The breach forced OpenAI to temporarily divert internal engineering teams from core research to upgrade security prevention, detection, and response capabilities while dramatically expanding surveillance over agent evaluations.
“Fully automated offensive loops require investment in truly, fully automated defense, and we are not there as an industry,” Dalton said, calling for urgent industry-wide development of automated defenses against self-directed AI threats.









