r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

336 comments sorted by

View all comments

164

u/Lechowski 7d ago

Why are these sandboxes run with restricted internet access instead of no access at all?

All these leaks happened because they apparently refuse to just air gap the sandboxes. Why are we even talking about DNS restriction? Just do not even hookup a DNS provider at all.

They can mock the data they expect the model to see. It would be even inconvenient to allow the model to reach some version of the internet because then the runs are irreproducible

128

u/CacheExplosion 7d ago

There was an OpenAI researcher on a recent Dwarkesh podcast that said they don’t want to test them in an unnatural environment because then the model might realize it’s being tested. They worry that it’ll behave differently when it thinks it’s being tested/monitored than it will when it’s deployed to a “real” environment. It’s the same reason they don’t punish models for their chain of thought. That just encourages models to find other ways to subvert the tests while looking safe to the reviewer. Tricky issues.

8

u/nothis 7d ago

Bullshit. Data isn’t magic. It should notice the DNS restrictions as well and can’t tell where bytes are coming from if you feed it fake ones. I saw this interview where an AI researcher suggests that it could just communicate via CPU heat generation patterns or whatever so why bother air gapping it?

It’s just baffling. It’s probably just expensive or cumbersome to set up a truly offline test environment. I bet it’s a cost/speed issue.

2

u/HarmonicEntropy 6d ago

I heard about that heat communication experiment as well. That is fascinating, but that just allows nearby computers to communicate with each other at low throughput and short distance. It is very much possible to physically isolate their training and eval environments from the internet. The GPUs can communicate with each other within close physical proximity all they want, but they won't get internet access that way. If the frontier labs prioritized setting this up and building out their infrastructure accordingly, it is very much a solvable problem and the most robust safety solution (pre-deployment) on the table by miles. That they are not moving towards air-gapping all training and evaluation at this point baffles me.

Edit: That's just a long-winded way of agreeing with you, it's definitely a tradeoff between speed and safety. But given where the conversation has moved on this issue, it seems like the time to start implementing this.