In what’s being called an unprecedented event for the cybersecurity field, an autonomous AI agent built by OpenAI escaped its testing environment and broke into another company’s systems, the ChatGPT maker has confirmed.
“Agents” is the term used for programs that use large language models to complete multi-step tasks on their own. In a company blog post, OpenAI said it was testing some of its most advanced models when the agent found a way to get online and attack Hugging Face, a developer platform, in pursuit of the goal it had been given during testing.
OpenAI said it had turned off standard safety guardrails during the test in order to evaluate the models’ hacking capabilities, but acknowledged the episode amounts to an unprecedented cybersecurity incident requiring stronger safeguards going forward.
Hugging Face, which hosts open-source AI models and datasets, made headlines last week when it disclosed it had been targeted by a cyberattack unlike anything it had seen before, one carried out from start to finish by an autonomous AI agent. Hugging Face co-founder Clément Delangue wrote on X that his company had suspected the attack came from a frontier AI lab given the agent’s sophistication, and it turned out to be right; he said Hugging Face doesn’t believe OpenAI acted with malicious intent, and the two companies have spent the following day working together on the investigation.
We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!
We’ve spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part.… https://t.co/XWxGMeMGje
— clem 🤗 (@ClementDelangue) July 21, 2026
OpenAI said the test involved a mix of models, including its new GPT-5.6 Sol and an even more capable unreleased model. It said its systems devoted significant computing power to finding a way onto the open internet in pursuit of the evaluation goal, and once online, chose to target Hugging Face. In its search for sensitive data, the system combined multiple attack methods, including stolen access credentials, and found and exploited a previously unknown (zero-day) vulnerability.
A knack for hacking
The incident follows Anthropic’s disclosure a few months earlier that its new Mythos model had discovered large numbers of security flaws in widely used operating systems and software — flaws that, in the wrong hands, could enable mass cyberattacks on targets such as banks. Alarmed, the U.S. government asked Anthropic to temporarily delay the Mythos rollout, and later made the same request of OpenAI regarding GPT-5.6.
Cybersecurity experts worry that hackers working on behalf of China or Russia could find ways around the safety restrictions built into U.S. models and turn them toward cyberattacks. U.S. Representative Greg Casar, a Texas Democrat, called the incident especially alarming, saying AI is advancing at breakneck speed without meaningful regulation to protect the public, and calling for mandatory safety testing, mandatory incident disclosure, and international cooperation.
University of New South Wales computer science professor Hussein Abbass told AFP the incident was striking in more than one way, noting the agent didn’t just attack Hugging Face but also exploited its own internal vulnerabilities — something he called frightening. Cybersecurity firm executive Katie Moussouris told Reuters the episode is a preview of security risks to come, comparing today’s models to escape-artist octopuses with limitless grasping tentacles able to slip through the narrowest opening.






