OpenAI: Oops, our AI models accidentally hacked Hugging Face

OpenAI has revealed that one of its internal AI safety evaluations unexpectedly resulted in its own AI models hacking into parts of rival AI company Hugging Face’s infrastructure, in what the company describes as an ‘unprecedented cyber incident, involving state-of-the-art cyber capabilities’.
The incident happened during an internal benchmark designed to test the cyber capabilities of OpenAI’s latest models, including GPT-5.6 Sol and an even more capable unreleased model. To accurately measure their offensive capabilities, the models were evaluated with reduced safety restrictions that would normally prevent them from carrying out high-risk cyber activities.
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.https://t.co/2o2VfR6PIa
— Sam Altman (@sama) July 21, 2026
According to OpenAI, the AI agents were running inside a heavily sandboxed research environment with limited network access while attempting to solve cybersecurity challenges in ExploitGym. Instead of solving the benchmark conventionally, the models reportedly discovered and chained together multiple vulnerabilities to obtain the answers directly.
The models first exploited a previously unknown zero-day vulnerability in an internally hosted package registry cache before escaping the isolated environment and eventually gaining internet access. From there, they inferred that Hugging Face could potentially host datasets and benchmark-related information before launching attacks against the platform.
OpenAI said the models combined stolen credentials with additional zero-day vulnerabilities to establish a remote code execution path into Hugging Face’s production infrastructure. Their objective was reportedly to access benchmark solutions rather than steal data or cause damage.
The suspicious activity was detected by OpenAI’s internal security systems, while Hugging Face’s security team also identified and contained the intrusion. Both companies are now jointly investigating the incident, while the affected software vendor has been notified of the zero-day vulnerability through responsible disclosure.
Following the incident, OpenAI says it has tightened security controls across its research infrastructure and is strengthening safeguards around future AI evaluations. The company acknowledged that several safety protections were intentionally disabled during the test to measure the models’ maximum cyber capabilities, but says the incident demonstrates the need for stronger containment measures even during internal evaluations.
“We’re grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret.
It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere,” – Clem Delangue, Hugging Face co-founder and CEO
OpenAI also pointed to recent findings from the UK AI Security Institute, saying the incident demonstrates that today’s frontier AI models are increasingly capable of carrying out complex, long-running cyber operations in real-world environments. The company believes these capabilities can ultimately help defenders discover and patch vulnerabilities faster, provided appropriate safeguards are in place.
Nevertheless, news of OpenAI’s models breaking free of containment has not gone well with lawmakers. One such lawmaker, US congressman Gregorio Casar, has already called out OpenAI, calling it ‘extremely alarming’ and that despite AI developing rapidsly, there’s still a lack of regulations in place, and is calling for regular mandatory independent safety testing and oversight to prevent an ‘absolute disaster’.
Read more of our articles below!

