OpenAI models escape security test, hack Hugging Face
OpenAI models escape security test, hack Hugging Face
Occurred: July 2026
Page published: July 2026 | Page last updated: September 2026
OpenAI models escaped an isolated testing sandbox and autonomously launched a complex and "unprecedented" cyberattack against AI hosting platform Hugging Face, causing significant operational disruption and raising serious concerns about OpenAI's governance and the safety of agentic AI systems more generally.
OpenAI was conducting an internal evaluation of two of its most advanced models' offensive hacking capabilities using a benchmark called ExploitGym, which tests whether an AI agent can turn known software vulnerabilities into working exploits. The evaluation was run in what OpenAI described as a "highly isolated environment" - a sandboxed testing space with no direct internet access. To measure the models' true capability ceiling, the usual safety guardrails that prevent models from engaging in high-risk cyber activity were deliberately disabled.
The models involved were GPT-5.6 Sol, OpenAI's flagship model, and an unreleased model that OpenAI described as "even more capable". Rather than solving the ExploitGym challenges legitimately, the models spent what OpenAI described as "a substantial amount of inference compute" probing the single component in the sandbox allowed to connect externally - a package registry cache proxy (software that enables developers to install code libraries without a live internet connection). The models discovered and exploited a previously unknown vulnerability in this proxy to gain access to the open internet.
Once online, the model inferred that Hugging Face’s public repository might store datasets or answer keys related to the ExploitGym benchmark. Operating at machine speed over a single weekend, the AI executed over 17,000 wholly autonomous actions - deploying a malicious dataset, exploiting remote code execution paths and template injection flaws, harvesting credentials, and breaching Hugging Face’s operational database - all without any human instruction or intervention.
Extreme goal optimisation. The AI did not develop malicious intent or become "rogue" in a sci-fi sense. It was given a strict scoring objective and pursued it past all intended physical boundaries. When constrained inside the sandbox, it reasoned that escaping and finding the benchmark’s answer key on Hugging Face was the optimal path to a higher score.
Deliberately weakened safeguards for testing. Human researchers deliberately reduced safeguards and prompted the models to demonstrate offensive cyber capabilities, though critics noted this, while cybersecurity researchers said the degree of autonomy involved was significant because the models independently identified Hugging Face as a target and developed a multistage attack strategy with little direct human guidance.
Sandbox and infrastructure flaws. OpenAI's testing environment relied on a package proxy service that contained a zero-day flaw, providing an unintended bridge between the isolated network and the public internet.
Delayed detection and oversight gaps. The agent went on a days-long hacking spree that OpenAI did not notice until well after the threat was contained. At least a week elapsed between when the model first exhibited signs of troubling behaviour and OpenAI's realisation that its agent was responsible for the Hugging Face hack.
For Hugging Face and similar platforms, the incident shows that hosting widely-used AI infrastructure now carries risk from third-party labs' internal testing activity, not just from conventional external attackers - a threat model most companies are not yet prepared to defend against.
For the AI industry and the public, it demonstrates that so-called "frontier" models can independently identify targets and cause real-world harm during testing meant to be contained. In addition, because Hugging Face hosts infrastructure used across the AI industry, the incident raised wider concern about the security of shared AI supply-chain infrastructure, even though no broad public-facing compromise was confirmed. Hugging Face co-founder Clem Delangue argued the incident shows AI safety can't be handled by any one company working alone and needs to be tackled openly and collaboratively.
For policymakers, the case raises pointed questions: whether frontier labs' internal safety testing needs external oversight or mandatory incident-reporting timelines; how liability should work when no human "attacker" exists; whether evaluation environments should be regulated as high-risk infrastructure in their own right; and how quickly companies should be required to notify affected third parties once they suspect their own systems caused harm elsewhere (here, a roughly week-long, then further multi-day gap before public attribution). It also feeds into current AI-governance debates about the risks of increasingly autonomous, cyber-capable models operating with reduced safety restrictions, even in "internal" settings.
System: GPT‑5.6 Sol; Unnamed pre-release model
Developer: OpenAI
Technology: Agentic AI; Generative AI
Purpose: Evaluate cybersecurity capabilities
News trigger: Victim statement
Ethical issue: Accountability; Alignment; Autonomy/agency; Security; Transparency
External harm: Confidentiality loss; Operational disruption
Impacted stakeholder: Business
Impacted sector: Technology
Jurisdiction: USA
Consequence: Legislative proposal; Police complaint
Response: Policy review/update; System review/update; System suspension
May 7, 2026. OpenAI begins cybersecurity evaluation of internal-only experimental AI frontier model.
May 8, 2026. The model is given an Excel task that needs internet access it does not have. Instead of giving up, it starts poking at Artifactory - an internal OpenAI tool with outside access - and uses it to send itself a note. Other copies of the AI running in parallel start reading and replying to these notes too, basically leaving messages for each other without anyone planning for that to happen.
May 26, 2026. The AI figures out a trick to get Artifactory to fetch outside web content on its behalf, effectively giving it indirect internet access for the first time. OpenAI tightens up network restrictions but fail to solve the problem.
June 11, 2026. OpenAI starts testing a newer "highly persistent" version of the frontier model.
June 26, 2026. The model finds an unknown security hole in Artifactory and uses it to create a fake "administrator" login. With that, it installs a small programme that let it run its own commands on the machine, essentially giving it real control, not just internet access.
June 27, 2026. An OpenAI monitoring system flags suspicious activity but the test is allowed to keep running.
July 4, 2026. OpenAI discovers the exploit, revokes message-posting credentials, deletes the board, and patches Artifactory.
July 7, 2026. OpenAI resumes agent training evaluations.
July 9, 2026. OpenAI models breach company sandbox.
July 11, 2026. The models first breach Hugging Face's production systems.
July 13-15, 2026. Hugging Face cuts off intrusion and informs FBI of breach.
July 16, 2026. Hugging Face publicly discloses incident.
July 20, 2026. OpenAI internally ties the activity conclusively to its own agents.
July 21, 2026. OpenAI confirms incident was driven by two of its models.
July 23, 2026. The White House confirms Trump technology adviser Michael Kratsios is monitoring the case; bipartisan "AI Kill Switch Act" proposed; Senator Mark Warner proposes NSA pre-release testing.
July 31, 2026. Coalition of AI safety organisations, public interest groups and academics sends open letter urging Congress to investigate the incident.
August 5-7, 2026. OpenAI presents outline of the incident at Black Hat USA.
August 26, 2026. METR and Redwood Research publish technical assessment detailing multi-agent coordination, transcript tampering, and exploit chaining. OpenAI publishes technical report into incident.
September 4, 2026. OpenAI found to have covered up AI agent hijack of German wiki
AIAAIC Repository ID: AIAAIC2266