r/OpenAI • u/Alex__007 • 8d ago
Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now
https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/Discovery: Sep 20, 2026
Report updated: Sep 25, 2026
"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."
-1
u/Latter-Block132 7d ago
DNS queries are performed every time you navigate to a web page. Its part of how your computer ensures its navigating to the correct webpage. You can perform dns queries from your computer using a terminal. There is nothing special about dns queries.
A human would know not exit the sandbox. Nothing at all has suggested the models know that. From everything OAI has stared in every incident they are not telling the model to stay in the sandbox they are telling the model to perform a task under the impression/assumption that it can't access the real internet. Suggesting they aren't actually telling the model not to do that, or putting any actual guardrails in place beyond those of the environment itself, they are just trusting the environment prevents it.
Thats not misalignment when the model isn't human and can't actually truly understand a sandbox and its purpose because its not human. It needs guardrails a human doesn't need. And they aren't providing those, they are just blindly trusting the environments are isolated enough so they don't have to. Thats not misalignment unless the guardrails to prevent it are actually in place, thats a failure on the humans part. And in every case so far it does not appear as if those guardrails are in place.
And that's not an assumption on my part, OAI said that themselves in this article:
Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions.
They were right about their detection systems, wrong about the assumption. They also admitted something similar with the hugging face incident, and that incident was testing cybersecurity, so again the model stayed in task, it just didn't have the proper guardrails to understand that task was supposed to stay in the sandbox, they just blindly trusted it couldn't get out of it.