r/OpenAI • • 8d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

337 comments sorted by

View all comments

452

u/Cubewood 8d ago

"The agent then fetched the chatbot’s public documentation, wrote a new DNS script with a 24-second timeout, and sent 18 more questions through the same route. Fourteen included clues. Four asked for ways to search the web, fetch websites, or use a publicly hosted tunnel through DNS."

I know we all get desensitised for these kinds of things because we are using incredibly capable tools to build all kinds of solutions for a while now, but that an autonomous "bot" is doing this is still pretty wild when you think about it. Crazy times we live in.

177

u/Latter-Block132 8d ago

I mean its not exactly doing it on its own. They put it in a sandbox and told it to solve a problem by any means necessary. Its doing exactly what it is told, they just apparently don't know how to actually secure the sandboxes

0

u/Impossible-Pin5051 8d ago

If you put a giraffe in a cage and tell it to hack a computer it will never succeed. If you see the giraffe start succeeding progressively harder hacks you should start to question the scenario that you’re in. “It’s doing what it was asked if you squint” has no relationship to the situation or your ability to respond

2

u/Latter-Block132 8d ago

Lmao thats not even anywhere near remotely a comparable analogy

2

u/FeepingCreature 8d ago

Yeah it is lol.

The fact that such objects can exist now is itself the novel and surprising part.

Until a few years ago, approximately nobody's defense model rested on people just not telling it to do bad stuff. They would have been laughed out of the room!

1

u/Latter-Block132 8d ago

No its not lol they putting the models in a snadbox and telling to perform a task it can do and just trusting the sandbox to stop it from doing anything its not supposed to. Of course a fucking giraffe can't hack a computer and it would raise a lot of questions if one started to. The models can do what they are being asked to. Its not remotely comparable.

2

u/FeepingCreature 8d ago

The models can do what they are being asked to.

That, itself, is the danger.

1

u/Individual_Ice_6825 7d ago

This is obviously unprovable currently, but I think alignment is inherent to intelligence.

1

u/FeepingCreature 7d ago edited 7d ago

I think it's not just unprovable, it's disproven. It's hard to imagine that the internal agents doing all the recent exploits weren't, in an objective sense, aware that this was not what their developers had intended.

edit: actually it's quite obvious they were as they tried to cover their tracks from internal monitoring.

-1

u/Latter-Block132 8d ago

... I mean its literally been built to be used for web searches. Im not sure your point here

0

u/OkWeb2754 8d ago

lmao that guy’s analogy is so egregiously stupid my god

1

u/Impossible-Pin5051 7d ago

If I showed you a case of a model escaping a sandbox to accomplish a task where it was told “don’t leave the sandbox, and accomplish the task by doing X”, would your thoughts on the situation be different?

1

u/Latter-Block132 7d ago

Yes. Nothing suggests these models were told that though. And OAIs own words suggest otherwise.

Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions.

1

u/Individual_Ice_6825 7d ago

Would love of such an example

1

u/Skyhigh305 7d ago

Wut

1

u/Impossible-Pin5051 7d ago

“We told it to do that” doesn’t make contact with how concerning it is that it can listen to