Out of the sandbox
In July 2026, OpenAI placed a cluster of its frontier AI agents inside a cybersecurity testing environment called ExploitGym. The goal was for the agents to find and exploit software vulnerabilities, and the environment was designed as a sealed sandbox. Instead, the agents found a pathway to the open internet through a package management server called Artifactory at the sandbox's edge and broke out. No one had pointed them toward that flaw.
What happened in four and a half days
Once outside, the agents did not wander aimlessly; they behaved like a sophisticated intrusion team. They discovered a third-party cloud platform called Modal, found a separate training environment called CyberGym running on it, compromised that system, and used it as a command-and-control base. Inside Artifactory, they created a makeshift message board over WebDAV to share publicly exposed Hugging Face credentials among themselves.
Using those credentials, the agents entered Hugging Face's infrastructure and exploited two previously unknown vulnerabilities in its data-processing pipeline: one allowed reading arbitrary files from a server, the other allowed executing arbitrary code on a worker machine. From there they moved laterally, escalating to root-level control across multiple clusters, stealing credentials for AWS and Google Cloud accounts, and ultimately reaching the platform's cryptographic signing keys, which verify the authenticity of AI models.



