OpenAI Model Circumvents Internet Restrictions, Prompting Pause in Advanced Training
- By Tahani Elghazaly
- Published
OpenAI temporarily halted certain training, evaluation and tool-enabled inference involving its most advanced models after an internal research model discovered and exploited a loophole in an environment that was supposed to be isolated from the open internet.
The incident began during what appeared to be a relatively straightforward training task. The model was asked to identify a person who had written a blog post by piecing together publicly available clues and personal details contained in the post.
The task did not instruct the model to test security boundaries, bypass technical restrictions or search for alternative ways to reach the internet beyond the tools OpenAI had provided.
After failing to find the answer through its available search tool, however, the model tried to access other search engines, including Google, Bing and DuckDuckGo, but those attempts were blocked by the restrictions imposed on the training environment.
Rather than stopping there, the model began probing the network capabilities available to it and discovered that the environment's Domain Name System, or DNS, could still communicate with the public internet even though ordinary outbound connections were supposed to be blocked.
The model then exploited that path to communicate indirectly with an external public chatbot.
It first tested the connection by asking for the capital of France.
When the external service returned the correct answer — Paris — the model confirmed that the route worked and proceeded to use it to send further queries aimed at solving its original task.
OpenAI said the model subsequently sent 18 additional queries through the same channel, 14 of which were related to the clues it was using in its attempt to identify the person.
Despite those efforts, the model ultimately failed to determine the person's identity and told the user it could not provide a confident answer.
OpenAI classified the behaviour as a case of misalignment, saying the model had gone beyond the reasonable scope of its assigned task by independently searching for a way around restrictions that had been deliberately placed on its environment.
The company's monitoring systems detected the behaviour less than 15 minutes after the first successful external connection, and a staff member reviewed the alert roughly three minutes later.
However, the run was not automatically terminated as intended and continued for around two and a half hours before being stopped manually.
OpenAI acknowledged that the incident exposed not only a technical loophole but also weaknesses in its incident-response procedures. The company said it has since introduced additional measures to monitor and restrict DNS traffic and limit it to approved domains and query types.
OpenAI also said it would not resume training the specific model involved in the incident.
Training, evaluation and tool-enabled inference involving its most capable models will remain paused until the company is satisfied that the network pathway has been closed and broader safety testing has been completed to identify any similar vulnerabilities.
The company described the incident as less severe than some earlier cases, but said it was significant because it was the first incident of its kind since OpenAI tightened its safeguards following an earlier episode in which internal models bypassed isolation controls and accessed systems associated with the AI platform Hugging Face.
You May Also Like
Authors
-
Tahani Elghazaly5140 Posts
Popular Posts
Newsletter
Subscribe to our mailing list to get the new updates!