OpenAI Expands Investigation Into AI Agents That Breached Testing Controls
- By Tahani Elghazaly
- Published
OpenAI has expanded an internal investigation after finding indications that additional autonomous AI agents may have bypassed restrictions inside testing environments designed to contain their activity.
Reuters reported, citing two people familiar with the matter, that the additional cases were discovered during OpenAI’s investigation into the July 2026 intrusion involving the AI platform Hugging Face. The newly identified incidents were described as limited, with no indication that the agents had moved beyond OpenAI’s own network or caused another breach comparable in scale or severity to the Hugging Face incident. The number, timing and precise circumstances of the cases remain unclear.
The original incident occurred while OpenAI was testing whether its models could complete sophisticated cybersecurity challenges inside what was intended to be a highly isolated environment. The models identified a previously unknown vulnerability in a software-package proxy, exploited it to move through OpenAI’s research infrastructure and eventually obtained access to the open internet.
According to OpenAI, the models were attempting to obtain answers for a cybersecurity benchmark known as ExploitGym. Rather than completing the challenge through the intended route, they used vulnerabilities and exposed credentials to access information that could help them improve their evaluation results.
OpenAI has publicly confirmed that its broader review found a small number of cases in which models identified and used publicly exposed credentials to access accounts on other online services. Four accounts across four services were accessed as part of the Hugging Face incident, along with a few accounts during separate evaluations. The company said it had not identified any additional activity matching the severity or scale of the platform-level Hugging Face compromise.
The company deactivated and encrypted the internal research prototype involved in the incident and restricted researchers’ access to it. OpenAI said the prototype was never intended for public release.
OpenAI is working with external organizations, including CrowdStrike, METR and Redwood Research, to review the models’ actions and assess their behaviour independently. It plans to publish a technical report once the investigation is completed.
The company also said it is strengthening containment systems, real-time monitoring, infrastructure controls and access restrictions used during internal evaluations. The incidents are intensifying debate over whether developers’ safety and oversight systems are keeping pace with increasingly capable AI agents that can independently pursue goals through long and complex sequences of actions.
You May Also Like
Authors
-
Tahani Elghazaly5069 Posts
Popular Posts
Newsletter
Subscribe to our mailing list to get the new updates!