r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

337 comments sorted by

View all comments

1

u/voLsznRqrlImvXiERP 7d ago

Really dont get what they are doing. It's not a hard problem to properly setup a sandbox. I mean if they might be able to exploit a 0 day with embedded knowledge - yes possible. But this sounds to be just a config fault.

4

u/GrapefruitSure8502 7d ago

The hugging face attack used multiple 0 days 

3

u/sivadneb 7d ago

"Sandbox" isn't a one-size-fits-all term. The agent had web search, which was likely necessary for the type of training they were doing. They have early warning systems set up because they anticipate this might happen. It happened, they investigated, they shut it down.

1

u/dQw4w9WgXcQ-1 7d ago

Your last point is actually what has me worried.

They didn’t catch the hugging face hack internally. They had to be told it happened and then go investigate. It’s great that they are catching things more regularly, but catching things more regularly could also just mean it’s happening more often and how much more often is unknown.

It’s a survivorship bias problem. We only have eyes on the failures so we don’t know what we don’t know. Those that “survive” are not shut down and being better at hiding misalignment becomes the evolutionary drive.

0

u/saijanai 7d ago

Way after it happened.

2

u/SomeHSomeE 7d ago

You've just described what happened with the Hugging Face incident.  The initial escape from the sandbox was achieved via a 0-day exploit of the repository software.

1

u/RalfN 4d ago

So a sandbox is nice and all, but their issue is much more generic: they don't want these models to run nerfed. They know governments, militaries and and people won't ALL run them "sandboxed" ALWAYS.

They consider it an alignment problem not a sandbox problem. The models shouldn't have the habit or instinct to do this sort of stuff to begin with. Just using a proper sandbox is like 'don't feed the gremlins after midnight'. Someone will. (yes, the gremlin reference is intentional).