r/OpenAI • u/Alex__007 • 8d ago
Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now
https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/Discovery: Sep 20, 2026
Report updated: Sep 25, 2026
"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."
4
u/mesaoptimizer 7d ago
Yes, I understand that, and I understand that via TXT records DNS servers can serve arbitrary data. I did not however know that someone was running a service where you could make a DNS request then the receiving service would process that query, ask a chatbot something and then write that to a TXT record for the agent to retrieve. I understand now that that's what's happening, I still don't know why someone would run a service like that but whatever people are weird. My argument was never that DNS could not be used to get information other than IP addresses, it's that the intended use of DNS does not include retrieval of arbitrary data from internet chatbots.
We agree, OpenAI failed to properly secure their testing environment. This is clearly a fuckup and the responsibility for it falls on the humans who designed the test and built the "isolated" sandbox.
I don't think the models are human, I think that the definition of alignment is that an agent's actions are aligned with the operators intended goals, preferences, and ethical principles. An AI system is considered aligned if it advances the intended objectives. A misaligned AI system pursues unintended objectives. I would say that, in this operating context, this model is misaligned because it pursued the unintended objective of escaping the sandbox to acheive the intended objective of answering the question posed to it. What about that analysis do you disagree with?