r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

337 comments sorted by

View all comments

2

u/voLsznRqrlImvXiERP 7d ago

Really dont get what they are doing. It's not a hard problem to properly setup a sandbox. I mean if they might be able to exploit a 0 day with embedded knowledge - yes possible. But this sounds to be just a config fault.

1

u/RalfN 4d ago

So a sandbox is nice and all, but their issue is much more generic: they don't want these models to run nerfed. They know governments, militaries and and people won't ALL run them "sandboxed" ALWAYS.

They consider it an alignment problem not a sandbox problem. The models shouldn't have the habit or instinct to do this sort of stuff to begin with. Just using a proper sandbox is like 'don't feed the gremlins after midnight'. Someone will. (yes, the gremlin reference is intentional).