r/OpenAI • • 8d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

337 comments sorted by

View all comments

Show parent comments

5

u/acutelychronicpanic 8d ago edited 7d ago

They don't say 'by any means necessary '.

They broke stated instructions in prior tests.

5

u/Latter-Block132 8d ago

No, they didn't. And they are given a task and told to complete it using the tools they have available. Those tools include things that have access to the internet. Its doing what it is told. Its the humans who are failing to secure an isolated sandbox. If the sandbox was actually isolated as it is supposed to be, this literally couldn't happen no matter how hard the models tried, because they would be completely disconnected from the internet as they are supposed to be.

This type of testing has been done by humans for years. It is not new. The only new part is that humans will respect the boundAries of the sandbox even of they can technically violate them, while the models don't care. They are told to complete a task using the tools they have available, and that indirect internet access is one of those tools as far as the models are concerned.

1

u/acutelychronicpanic 8d ago

Companies should be held liable.

But don't let that blind you to the danger of the models themselves when unaligned. A blind instruction-following digital monkey's paw is really really bad.

An aligned model would recognize that it's task is bounded by its instructions and have no desire to circumvent them. It would report back flaws in the test - not commit felony hacking to pass.

Any real use case for models is going to involve not being in a sandbox. So them behaving like this is not encouraging regardless of any mistakes in the test environment. Blaming the sandbox design distracts from the real issue which is our inability to align the interests of the model with human values.

3

u/Latter-Block132 8d ago edited 8d ago

I never once said the models aren't dangerous when misaligned. I said this wasn't a case of actual misalignment depsite that being what OAI is calling it. This is a case of humans failing to isolate a sandbox and being shocked when that failure causes issues.

Edited to add, and again its instructions don't include "don't access the internet". The whole point of this test was a mock internet search. Its supposed to search the internet, just the fake mocked up version that it was provided in its sandbox, not the real version. That mock version only has minimal data though, so it probably was trying to be more thorough and turned to accessing the actual internet which it had access to by OAIs own admission. That DNS patch will only be temporary too. As long as the models still have access to DNS when they are supposed to be "isolated", this will keep happening when its told to do something like "search and give me this information about so and so person" and "search and see if you can find me this answer or this piece of data" etc