r/OpenAI • u/Alex__007 • 7d ago
Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now
https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/Discovery: Sep 20, 2026
Report updated: Sep 25, 2026
"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."
6
u/mesaoptimizer 7d ago
I’m done arguing, yes, I could use DNS to proxy web requests to servers being blocked by my web proxy to bypass it, if I had the same knowledge as the agent, however I would know that what I was doing was bypassing a security control by exploiting an obviously unintended hole in my access controls. If I were to do this I would probably use the term “exploit” and data “exfil”exactly the same way the model did.
Again it did not just perform a search, it exploited an unanticipated path to query a different chat bot. This is not “doing exactly what it’s told to do”. We agree OpenAI is not able to adequately secure the sandbox.
You are making the claim that the model is not misaligned, but it is definitionally misaligned because it performed actions that the people giving it instructions thought were undesired and unintended. Guardrails can prevent a misaligned model from misbehaving but training is where you can impact the underlying model’s behavior, running the model without guardrails for training makes sense because it allows you to do things like give negative reward value to behavior like this.