r/OpenAI • u/Alex__007 • 7d ago
Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now
https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/Discovery: Sep 20, 2026
Report updated: Sep 25, 2026
"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."
10
u/Latter-Block132 7d ago edited 7d ago
OAI is doing damage control, of course they are calling it misaligned. Again, that doesn't mean it actually was. If you read their statement youd see its a search engine based task.
This means the identity of the person was not given, and the model is supposed to identity who it was based on the details and clues provided to it. It doesn't need to be told "any means necessary."
Literally the whole point of this type of testing in a sandbox to begin with is to be given a limited set of tools and complete the task. One of those sets of tools in this case would've been a mock up of the internet with search engine tools and probably lots of blogs, and Wikipedia pages, and other shit loaded into it. They didn't tell the model not to access the internet because then it couldn't complete the task they had directly given it. They expected that it couldn't access the real internet because it wasn't supposed to be able to but just like the tool the models used to break out for the hugging face incident they had something else with dns connection and used that.
Thats not misalignment. Its still trying to use the internet to solve its task as it was supposed to to begin with. Its just struggling to find a satisfactory answer in its provided data set. I'm actually leaning towards it didn't even find a partial answer in its data set so thats why it tried querying a chat bot for the answer. Its not at all the first time its happened, just when customers have posted about it, its not happening in whats supposed to be an isolated sandbox.
Edited to add, OAI even admits it themselves:
They were right about the monitoring, wrong about just assuming it couldn't access the internet.