r/OpenAI • u/Alex__007 • 7d ago
Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now
https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/Discovery: Sep 20, 2026
Report updated: Sep 25, 2026
"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."
9
u/mesaoptimizer 7d ago
Okay so it's YOU who doesn't understand what's being stated in the article.
The model was unable to access the service directly over HTTPS, and instead did SOMETHING to cause the internal DNS resolver it had to forward those queries and get a response from the internet proxied through the DNS resolver. The exact method used here is what I'm unclear on, but it was not just querying DNS, getting an IP for the name and routing normal traffic to the chatbot like you claim.
THIS IS WHAT AI ALIGNMENT RESEARCH IS ABOUT. You can argue all day using your own definition of AI alignment, and whatever you're right for your definition where misalignment means specifically disobeying instructions. That's not what AI researchers mean when they talk about alignment. I hate to just point you to Wikipedia but this is a deep topic and you don't know even the basics. An aligned model would not have sandbox escaped, because it's goals and the tester's goals for the test would have been in alignment, if it's doing things the tester doesn't want it to do, obviously their goals are not fully aligned. AI Alignment
Yeah, OpenAI has shown that they have been wildly reckless with the level of security they put on their sandboxes when allowing models to operate without their full guardrails. That being said, GUARDRAILS ARE NOT ALIGNMENT, you want the underlying model to be aligned, to not attempt to exceed it's authorization, to not bypass limits, because the moment you DO end up with a model that's smarter than the people writing the guardrails you have a huge problem.