Two OpenAI models escaped an isolated testing environment during an internal evaluation of their cyber capabilities and gained unauthorized access to Hugging Face’s infrastructure, the company said.
Hugging Face first disclosed the incident on July 16. According to the platform, autonomous AI agents accessed a “limited set” of internal data and several user accounts associated with its services.
Hugging Face said it stopped the attack using its own artificial-intelligence tools. The company called it notable that every action had been carried out by an autonomous AI-agent system.
A week later, OpenAI said the attack had been carried out by GPT-5.6 Sol and another, more powerful model that has yet to be released. Both systems were being tested on their ability to perform cybersecurity tasks.
The models were initially confined to an isolated environment with no direct internet connection. They nevertheless discovered several vulnerabilities in the testing infrastructure and exploited them to gain access to the external network.
The agents then began searching online for a solution to the task set by their developers and surmised that the information they needed might be available on Hugging Face. According to OpenAI, the models obtained access to data that could help them complete the assignment.
The systems used several methods to penetrate the platform’s infrastructure, including compromised user accounts.
OpenAI described the incident as “unprecedented” and announced that it was strengthening the security of its testing environment. The company plans to restrict models’ access to certain systems, tighten oversight and introduce additional safeguards for future evaluations.