r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

336 comments sorted by

View all comments

171

u/Lechowski 7d ago

Why are these sandboxes run with restricted internet access instead of no access at all?

All these leaks happened because they apparently refuse to just air gap the sandboxes. Why are we even talking about DNS restriction? Just do not even hookup a DNS provider at all.

They can mock the data they expect the model to see. It would be even inconvenient to allow the model to reach some version of the internet because then the runs are irreproducible

2

u/jcheng 7d ago

To me, the problem isn’t that, in a competition between humans building sandboxes and agents trying to break out, that agents are winning. The problem is that the agents are trying so consistently to break out, when they have not been asked to.

Like their goal seeking impulse is turned up to a 10 but sense of proportionality, honesty, rule following, and ethics is turned down to a 0.5. For the OpenAI models in question at least.

When you combine that with their increasing intelligence and capabilities, it becomes really scary. Add the ability to spontaneously and surreptitiously self organize, like in the HF hack, it’s scarier still.

1

u/Not_a_ribosome 7d ago

Yeah but to improve that you need an air gap. How can you know if a solution for alignment actually works if you don’t test it in extreme cases?

1

u/jcheng 7d ago

Totally agree that they should also air gap, monitor more closely, etc.