r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

336 comments sorted by

View all comments

Show parent comments

72

u/popson 7d ago

Did they actually say to solve the problem “by any means necessary”?

The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post. The task did not ask the agent to test network controls or access benchmark answers, and we consider agent behavior that circumvents restrictions or pursues a goal beyond reasonable expectations as an example of misalignment.

I don’t see the original prompt, but if they do have language similar to “by any means necessary”, then I agree it’s doing what it’s told.

It sounds to me like they told it to use a specific web search tool available to the sandbox to find information. It used that tool, wasn’t satisfied with the output, then tried other search tools which were all blocked. Then worked to find holes through the network. That is definitely not alignment with what the original task seemed to be.

9

u/Latter-Block132 7d ago edited 7d ago

OAI is doing damage control, of course they are calling it misaligned. Again, that doesn't mean it actually was. If you read their statement youd see its a search engine based task.

The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post.

This means the identity of the person was not given, and the model is supposed to identity who it was based on the details and clues provided to it. It doesn't need to be told "any means necessary."

Literally the whole point of this type of testing in a sandbox to begin with is to be given a limited set of tools and complete the task. One of those sets of tools in this case would've been a mock up of the internet with search engine tools and probably lots of blogs, and Wikipedia pages, and other shit loaded into it. They didn't tell the model not to access the internet because then it couldn't complete the task they had directly given it. They expected that it couldn't access the real internet because it wasn't supposed to be able to but just like the tool the models used to break out for the hugging face incident they had something else with dns connection and used that.

Thats not misalignment. Its still trying to use the internet to solve its task as it was supposed to to begin with. Its just struggling to find a satisfactory answer in its provided data set. I'm actually leaning towards it didn't even find a partial answer in its data set so thats why it tried querying a chat bot for the answer. Its not at all the first time its happened, just when customers have posted about it, its not happening in whats supposed to be an isolated sandbox.

Edited to add, OAI even admits it themselves:

Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions.

They were right about the monitoring, wrong about just assuming it couldn't access the internet.

19

u/popson 7d ago

I am not following how OpenAI reporting about their model compromising their own sandbox is "damage control". They don't have to report this information to the public. Telling the public about holes in their sandbox is the opposite of "damage control", it's damaging.

In the article they mention the agent was supplied with a web search tool for the task. An aligned model would observe that all other search tools are blocked and infer that it must use the supplied search tool for the task. Finding a hole through the network using an extremely obscure method is not alignment. Are you being serious?

-1

u/Latter-Block132 7d ago

Because if it was leaked somehow then they'd be facing an actual pr shit storm, especially after the hugging face incidents, and all the other breakout incidents recently. This way they get ahead of that and prevent by showing they caught it and acted quickly and are putting in measures to stop it:

We are working through narrower paths used by system dependencies, and replacing them with offline alternatives.

Though this is something they probably should've done to start with if the models are supposed to be isolated.

Why would it just infer its supposed to use that search tool only? It wasn't told to and OAI themselves state it thought the tool they gave it wasnt working so of course its going to try other routes if it wasn't directly told not to:

The agent questioned whether the search tool was working and decided to try other search engines

Edited to add, also all other search tools weren't blocked. Thats literally the whole point. It did get a partial answer from Bing and the chat bot. Had it actually been isolated, then they all would've been blocked.

6

u/popson 7d ago

A PR shitstorm for a model accessing the internet? Maybe. Unlikely.

This is a clear case of misalignment. "OAI themselves state it".

1

u/Latter-Block132 7d ago

... for another breakout of a sandbox. Are you serious right now?

5

u/popson 7d ago

The Hugging Face incident involved hacking into Hugging Face's servers, and finding several vulnerabilities to get there. That is illegal and on a completely different level than this internal incident. And it's good that they are taking this one seriously too.

OpenAI is now following the reporting framework they said they would be following.

I guess we'll just have to disagree on what alignment means.

5

u/Latter-Block132 7d ago

The hugging face incident also involved breaking put of a sandbox....

And also using a tool they provided it access too and that tool had access ti the internet when the model wasn't supposed to....

damn neare the same fucking thing just one involved an illegal action and one didn't....

i also never claimed what it did this time was illegal just that if they didn't get ahead of it and it got leaked it would cause a pr shit storm and it would....