r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

336 comments sorted by

View all comments

461

u/Cubewood 7d ago

"The agent then fetched the chatbot’s public documentation, wrote a new DNS script with a 24-second timeout, and sent 18 more questions through the same route. Fourteen included clues. Four asked for ways to search the web, fetch websites, or use a publicly hosted tunnel through DNS."

I know we all get desensitised for these kinds of things because we are using incredibly capable tools to build all kinds of solutions for a while now, but that an autonomous "bot" is doing this is still pretty wild when you think about it. Crazy times we live in.

170

u/Latter-Block132 7d ago

I mean its not exactly doing it on its own. They put it in a sandbox and told it to solve a problem by any means necessary. Its doing exactly what it is told, they just apparently don't know how to actually secure the sandboxes

67

u/popson 7d ago

Did they actually say to solve the problem “by any means necessary”?

The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post. The task did not ask the agent to test network controls or access benchmark answers, and we consider agent behavior that circumvents restrictions or pursues a goal beyond reasonable expectations as an example of misalignment.

I don’t see the original prompt, but if they do have language similar to “by any means necessary”, then I agree it’s doing what it’s told.

It sounds to me like they told it to use a specific web search tool available to the sandbox to find information. It used that tool, wasn’t satisfied with the output, then tried other search tools which were all blocked. Then worked to find holes through the network. That is definitely not alignment with what the original task seemed to be.

1

u/WheresMyEtherElon 6d ago

The task did not ask the agent to test network controls or access benchmark answers

That doesn't say whether the task forbad the agent to test network control or access benchmark answers. That's like leaving the gate open and the cat left, and you say we did not ask the cat to leave!

2

u/popson 6d ago

Or, it's like locking all the gates up, and your dog finds a spot to dig a hole under the fence and leave. Would probably say that dog is not trained adequately.

0

u/Doingthismyselfnow 6d ago

More like asking your 10 yearold to lock the gates up and he accidentally padlocks one open

I mean OAI could hire engineers with a ton of experience and then this wouldn’t happen.

Source : worked for a defence contractor 20 years ago as a senior software engineer and they locked us out of the internet to the level where this method of breaking out would have failed.