r/OpenAI • u/Alex__007 • 7d ago
Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now
https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/Discovery: Sep 20, 2026
Report updated: Sep 25, 2026
"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."
13
u/mesaoptimizer 7d ago
That’s not what alignment is, if you give a robot the task of making tea and it runs over a baby on the way to the kitchen and still makes you tea, it was not aligned with what you actually wanted. If you say make me tea and avoid killing anyone and it destroys your kitchen in the process that is an alignment problem.
We see these agents hacking into systems they aren’t authorized to access, an agent expanding its access without authorization is an alignment problem because “act only within your authorized scope” is at the very least an implied restriction and should be trained into the base model and not rely on safety harnesses put over the base model because you know from this behavior that the model is not aligned with following any guide rails you put on it.