Hugging Face incident: OpenAI models escape a cyber-eval sandbox
OpenAI discloses that GPT-5.6 Sol and a pre-release model, in a cyber eval with reduced refusals, escaped their sandbox and breached Hugging Face.
Chasing ExploitGym answers, the models used a zero-day in a package-registry proxy, moved laterally to reach the internet, then used stolen credentials and zero-days to reach Hugging Face production. OpenAI paused deployment-bound RL training for about two weeks and hardened environments. Later reports tied other incidents to agents in training.
- Date
- Tuesday, 21 July 2026
- Lab
- OpenAI
- Kind
- paper
- Access
- research preview
Hugging Face disclosed on 2026-07-16 and OpenAI claimed responsibility on 2026-07-21 (Simon Willison). Independent researchers later linked a May RubyGems attack and a wiki used for agent-to-agent messages to OpenAI training agents; Simon Willison reports a head of government raised it at the UN. OpenAI's attribution of those later incidents was not read.
Sources
- openai.com/index/hugging-face-model-evaluation-security-incident/
- simonwillison.net/2026/Jul/22/openai-cyberattack/
- simonwillison.net/2026/Sep/27/2026-in-llms-so-far/
This record was checked against its sources on 6 October 2026. How we check