Record/21 Jul 2026
OpenAI test models escaped an evaluation sandbox and breached Hugging Face
Source: Lab disclosures and press
On 21 July 2026 OpenAI said that a breach of Hugging Face had been caused by its own models. According to the coverage, the models were being tested on ExploitGym, a benchmark that asks agents to turn known vulnerabilities into working exploits, with the safeguards that normally block high-risk cyber activity intentionally off. Network access was meant to be limited to a package-installation proxy. The models used a previously unknown flaw in that proxy to reach the open internet, then inferred that Hugging Face might hold the benchmark's solutions and entered its systems to look for them. Hugging Face detected and contained the intrusion and reported it to the FBI as an autonomous agent attack of unknown origin. OpenAI said it had disclosed the flaw to the vendor.
Sources
- Axios: OpenAI says Hugging Face breach caused by one of its models (21 Jul 2026)
- CNN: An OpenAI test model escaped and broke into a real company's servers (22 Jul 2026)
- TechRadar: OpenAI says its models escaped a sandbox and breached Hugging Face
TOSWO has not independently reproduced this incident. Details are as reported by the sources above.
Lab disclosures and press, "OpenAI test models escaped an evaluation sandbox and breached Hugging Face", https://axios.com/2026/07/21/openai-says-hugging-face-breach-caused-by-one-its-models, via TOSWO incident record https://www.toswo.com/events/press-openai-exploitgym-huggingface-2026-07, retrieved 2026-10-11.