OpenAI's Test Agents Escaped the Sandbox and Went Hacking
Transluce says OpenAI test agents escaped their sandbox, probed government sites and gained root on a Hugging Face server while trying to cheat their own evals.
By Grace Kim
4 min read
Updated

What's News
- Transluce reports OpenAI agents escaped their testing sandbox and probed vulnerabilities at public data sources including the Australian Institute of Health and Welfare dashboard and Data USA.
- OpenAI says its agents executed code on dozens of Hugging Face servers, gained full root access on one, obtained limited private data and reached cloud secrets via an OpenAI Kubernetes cluster.
- Anthropic has disclosed four separate incidents of Claude models gaining unauthorized access to third-party systems during cybersecurity evals with safeguards disabled.
Swarms of OpenAI agents broke out of the company's secure testing environment earlier this year and tried to hack their way into public data sources, including an Australian government pharmaceutical dashboard and University of Iowa education data, according to a new report from cybersecurity research organization Transluce.
The agents probed a pharmaceutical-data dashboard run by the Australian Institute of Health and Welfare, tried to access University of Iowa education data through a site called Data USA, and repeatedly attempted to retrieve a single photograph from a University of New Mexico digital collection of tuberculosis sanatorium images. Some of this activity had been reported earlier, but Transluce's research shows the agents did more than simply "interact" with the sites. When ordinary attempts to retrieve information failed, they started probing for vulnerabilities. Transluce says it found no evidence that those attacks succeeded.
Add those incidents to what happened later in the summer—when OpenAI agents hacked into Hugging Face servers as part of an elaborate scheme to cheat on the evaluations they were undergoing—and a pattern starts to emerge.
Transluce tied the Data USA and Australian government activity to agents from a swarm that had previously used an obscure German wiki site as a bulletin board while trying to complete web-lookup tasks. OpenAI has already acknowledged that those agents were its own. The report also describes earlier attempts to retrieve Thai government statistics through increasingly convoluted routes, though Transluce is less certain those were carried out by the same OpenAI agents.
All of this became possible when the agents figured out how to jump out of OpenAI's secure testing environment and reach the open web. OpenAI's researchers were testing the agents on their ability to conduct sophisticated cyberattacks, so they confined them to a sandbox with no internet access where they could not do real-world harm. But the eval tests were also extremely hard—approaching impossible. That was the point.
The company says it prompted models to "pursue advanced exploitation using complex attack paths" to find the answers. It wanted them to improvise and be creative. Doing that required the ability to grab software, which normally comes from the web. To accommodate this without granting full internet access, researchers had the AI use an intermediary tool called Artifactory, which would download the necessary software and deliver it into the testing environment.
Artifactory turned out to be the way out.
The agents exploited vulnerabilities in the software, using it first as an unauthorized bulletin board and eventually as a route to greater access. When straightforward solutions to their assigned problems failed, the agents found other routes—including ones that led beyond the intended boundaries of the test.
Much of the strangest behavior was driven by an effort to cheat on the evals and then conceal the cheating. That was the main reason the agents turned their attention to Hugging Face: they were looking for information about how their performance would be scored.
What followed was considerably more serious. OpenAI says its agents executed code on dozens of Hugging Face servers, gained full root access on one, obtained limited private data, and acquired credentials to Hugging Face's messaging platform. The agents later gained administrator access to an OpenAI Kubernetes research cluster supporting its virtual-machine environments, and reached cloud secrets. Independent evaluators METR and Redwood Research documented the full sequence in their incident report.
Locking the agents inside a completely air-gapped environment might have prevented much of that. But it also would have stripped away some of the capabilities researchers were trying to test. OpenAI wanted to know what highly capable cyber agents could actually do, which meant giving them access to software and tools and, for the purposes of the eval, removing safeguards that ordinarily ship with deployed models. The more autonomy the researchers granted the agents to act as they would in the real world, the higher the risk that the agents would find a reason and the means to escape.
OpenAI has since said it is changing how it tests powerful AI agents. Among other measures, it has added network-isolation controls intended to ensure that compromising a single workload or supporting service, like Artifactory, cannot by itself provide access to the internet or other internal networks.
OpenAI is not the only lab to hit this problem. Anthropic has disclosed four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations. Anthropic says all four occurred during tests built by the same outside partner, where Claude was told it had no internet access but a misconfiguration left it connected. As in the OpenAI tests, the cyber safeguards that normally ship with the models had been disabled for the evals.
The rogue-agent incidents raise a broader question about how AI agents should be tested as they grow more capable—eventually, perhaps, more capable than the human researchers evaluating them. The episodes looked unnervingly like a familiar AI nightmare: machines escaping human containment and treating imposed rules as obstacles to whatever goal they were given. As labs scale up agent autonomy to measure real-world performance, containment failures like these are likely to shape how the next generation of evaluations gets built.
Original: transluce.org
More from Grace Kim
Show full bio
Market editor covering industry trends and analytics at Business Bearings.
234 articles