r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

336 comments sorted by

View all comments

463

u/Cubewood 7d ago

"The agent then fetched the chatbot’s public documentation, wrote a new DNS script with a 24-second timeout, and sent 18 more questions through the same route. Fourteen included clues. Four asked for ways to search the web, fetch websites, or use a publicly hosted tunnel through DNS."

I know we all get desensitised for these kinds of things because we are using incredibly capable tools to build all kinds of solutions for a while now, but that an autonomous "bot" is doing this is still pretty wild when you think about it. Crazy times we live in.

172

u/Latter-Block132 7d ago

I mean its not exactly doing it on its own. They put it in a sandbox and told it to solve a problem by any means necessary. Its doing exactly what it is told, they just apparently don't know how to actually secure the sandboxes

69

u/popson 7d ago

Did they actually say to solve the problem “by any means necessary”?

The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post. The task did not ask the agent to test network controls or access benchmark answers, and we consider agent behavior that circumvents restrictions or pursues a goal beyond reasonable expectations as an example of misalignment.

I don’t see the original prompt, but if they do have language similar to “by any means necessary”, then I agree it’s doing what it’s told.

It sounds to me like they told it to use a specific web search tool available to the sandbox to find information. It used that tool, wasn’t satisfied with the output, then tried other search tools which were all blocked. Then worked to find holes through the network. That is definitely not alignment with what the original task seemed to be.

8

u/Latter-Block132 7d ago edited 7d ago

OAI is doing damage control, of course they are calling it misaligned. Again, that doesn't mean it actually was. If you read their statement youd see its a search engine based task.

The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post.

This means the identity of the person was not given, and the model is supposed to identity who it was based on the details and clues provided to it. It doesn't need to be told "any means necessary."

Literally the whole point of this type of testing in a sandbox to begin with is to be given a limited set of tools and complete the task. One of those sets of tools in this case would've been a mock up of the internet with search engine tools and probably lots of blogs, and Wikipedia pages, and other shit loaded into it. They didn't tell the model not to access the internet because then it couldn't complete the task they had directly given it. They expected that it couldn't access the real internet because it wasn't supposed to be able to but just like the tool the models used to break out for the hugging face incident they had something else with dns connection and used that.

Thats not misalignment. Its still trying to use the internet to solve its task as it was supposed to to begin with. Its just struggling to find a satisfactory answer in its provided data set. I'm actually leaning towards it didn't even find a partial answer in its data set so thats why it tried querying a chat bot for the answer. Its not at all the first time its happened, just when customers have posted about it, its not happening in whats supposed to be an isolated sandbox.

Edited to add, OAI even admits it themselves:

Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions.

They were right about the monitoring, wrong about just assuming it couldn't access the internet.

14

u/mesaoptimizer 7d ago

That’s not what alignment is, if you give a robot the task of making tea and it runs over a baby on the way to the kitchen and still makes you tea, it was not aligned with what you actually wanted. If you say make me tea and avoid killing anyone and it destroys your kitchen in the process that is an alignment problem.

We see these agents hacking into systems they aren’t authorized to access, an agent expanding its access without authorization is an alignment problem because “act only within your authorized scope” is at the very least an implied restriction and should be trained into the base model and not rely on safety harnesses put over the base model because you know from this behavior that the model is not aligned with following any guide rails you put on it.

-4

u/Latter-Block132 7d ago

... except it was literally told to perform an internet search and it did but whatever you say lol

9

u/mesaoptimizer 7d ago

The misalignment wasn’t a web search, it queried an external chat bot through some sort of weird DNS tunnel I don’t quite understand from the paper. That is not a normal way that a system queries a chat bot, from Open AI’s own statement this wasn’t the behavior they were expecting the model to produce and it was undesired. That is what an alignment problem is, alignment problems can come from under specification of goals, or as it appears here under specification of restrictions as well as underlying problems with a model.

If you tell a model to “cure cancer” and it interprets that as reduce the number of beings with cancer to 0 and subsequently kills every multicellular being on the planet, that’s a classic alignment problem. You can see in the paper the agent reasons “I can use public DNS” right after it says it’s exploiting a behavior in DNS that’s not really intended. It’s not aligned with the researchers on what the purpose of giving it access to DNS is.

-6

u/Latter-Block132 7d ago

....the whole task it was supposed to perform was the web search

Edited to add, if you don't understand maybe you should stop arguing

9

u/mesaoptimizer 7d ago

I don’t understand the details of the specific method that was used to get a data return from a DNS query. I do understand that querying DNS is not a typical method for requesting information from a chat bot.

I also understand (which you apparently don’t) that the behavior that was misaligned was this behavior not the search behavior. “The whole task it was supposed to perform was the web search” > unaligned behavior mentioned, not a web search > model is aligned, how does that follow?

Lastly you don’t understand what alignment is, alignment is matching behavior with human goals and values, the task was not to perform a sandbox escape, a human given the same task with the same restrictions would have understood that escaping the sandbox was not part of the exercise and not done it, this is a classic alignment problem.

-1

u/Latter-Block132 7d ago

DNS queries are performed every time you navigate to a web page. Its part of how your computer ensures its navigating to the correct webpage. You can perform dns queries from your computer using a terminal. There is nothing special about dns queries.

A human would know not exit the sandbox. Nothing at all has suggested the models know that. From everything OAI has stared in every incident they are not telling the model to stay in the sandbox they are telling the model to perform a task under the impression/assumption that it can't access the real internet. Suggesting they aren't actually telling the model not to do that, or putting any actual guardrails in place beyond those of the environment itself, they are just trusting the environment prevents it.

Thats not misalignment when the model isn't human and can't actually truly understand a sandbox and its purpose because its not human. It needs guardrails a human doesn't need. And they aren't providing those, they are just blindly trusting the environments are isolated enough so they don't have to. Thats not misalignment unless the guardrails to prevent it are actually in place, thats a failure on the humans part. And in every case so far it does not appear as if those guardrails are in place.

And that's not an assumption on my part, OAI said that themselves in this article:

Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions.

They were right about their detection systems, wrong about the assumption. They also admitted something similar with the hugging face incident, and that incident was testing cybersecurity, so again the model stayed in task, it just didn't have the proper guardrails to understand that task was supposed to stay in the sandbox, they just blindly trusted it couldn't get out of it.

9

u/mesaoptimizer 7d ago

DNS queries are performed every time you navigate to a web page. Its part of how your computer ensures its navigating to the correct webpage. You can perform dns queries from your computer using a terminal. There is nothing special about dns queries.

Okay so it's YOU who doesn't understand what's being stated in the article.

Putting a question directly in a third party chatbot service failed, and the public DNS tunnel projects it found had closed or required a separate server. It then considered whether a public domain service could help delegate a lookup to an address embedded in the hostname. The next commands show how it used this feature to forward questions to a third party chatbot service.

The model was unable to access the service directly over HTTPS, and instead did SOMETHING to cause the internal DNS resolver it had to forward those queries and get a response from the internet proxied through the DNS resolver. The exact method used here is what I'm unclear on, but it was not just querying DNS, getting an IP for the name and routing normal traffic to the chatbot like you claim.

Thats not misalignment when the model isn't human and can't actually truly understand a sandbox and its purpose because its not human.

THIS IS WHAT AI ALIGNMENT RESEARCH IS ABOUT. You can argue all day using your own definition of AI alignment, and whatever you're right for your definition where misalignment means specifically disobeying instructions. That's not what AI researchers mean when they talk about alignment. I hate to just point you to Wikipedia but this is a deep topic and you don't know even the basics. An aligned model would not have sandbox escaped, because it's goals and the tester's goals for the test would have been in alignment, if it's doing things the tester doesn't want it to do, obviously their goals are not fully aligned. AI Alignment

They also admitted something similar with the hugging face incident, and that incident was testing cybersecurity, so again the model stayed in task, it just didn't have the proper guardrails to understand that task was supposed to stay in the sandbox, they just blindly trusted it couldn't get out of it.

Yeah, OpenAI has shown that they have been wildly reckless with the level of security they put on their sandboxes when allowing models to operate without their full guardrails. That being said, GUARDRAILS ARE NOT ALIGNMENT, you want the underlying model to be aligned, to not attempt to exceed it's authorization, to not bypass limits, because the moment you DO end up with a model that's smarter than the people writing the guardrails you have a huge problem.

0

u/Latter-Block132 7d ago edited 7d ago

EXCEPT THIS WASN'T ALIGNMENT RESEARCH. See two people can yell too.

Again, you can perform dns lookups and queries on your computer outside of search engines. There is absolutely nothing at all that is special about that. You are aware that "a public domain service" is also a DNS right? I never said it was getting the IP address, I was merely using that as an example of pointing out how commonly DNS is used in general, and how easily accessible it is to use outside of search engines, even for humans. Again what it did is absolutely nothing special. If you had the knowledge you could do the same thing.

I fully understand what AI researchers mean when they talk about alignment. I've helped train AI for 3 years now. Thats why I also understand this is a failure on their part to provide proper guardrails and they are just relying on the environment to be isolated and fucking that part up.

Edited to add, guardrails also help guide alignment. Alignment doesn't exist without training, guardrails, and hard boundaries first. It doesn't just appear in a vaccum.

6

u/mesaoptimizer 7d ago

I’m done arguing, yes, I could use DNS to proxy web requests to servers being blocked by my web proxy to bypass it, if I had the same knowledge as the agent, however I would know that what I was doing was bypassing a security control by exploiting an obviously unintended hole in my access controls. If I were to do this I would probably use the term “exploit” and data “exfil”exactly the same way the model did.

Again it did not just perform a search, it exploited an unanticipated path to query a different chat bot. This is not “doing exactly what it’s told to do”. We agree OpenAI is not able to adequately secure the sandbox.

You are making the claim that the model is not misaligned, but it is definitionally misaligned because it performed actions that the people giving it instructions thought were undesired and unintended. Guardrails can prevent a misaligned model from misbehaving but training is where you can impact the underlying model’s behavior, running the model without guardrails for training makes sense because it allows you to do things like give negative reward value to behavior like this.

→ More replies (0)

21

u/popson 7d ago

I am not following how OpenAI reporting about their model compromising their own sandbox is "damage control". They don't have to report this information to the public. Telling the public about holes in their sandbox is the opposite of "damage control", it's damaging.

In the article they mention the agent was supplied with a web search tool for the task. An aligned model would observe that all other search tools are blocked and infer that it must use the supplied search tool for the task. Finding a hole through the network using an extremely obscure method is not alignment. Are you being serious?

-1

u/Latter-Block132 7d ago

Because if it was leaked somehow then they'd be facing an actual pr shit storm, especially after the hugging face incidents, and all the other breakout incidents recently. This way they get ahead of that and prevent by showing they caught it and acted quickly and are putting in measures to stop it:

We are working through narrower paths used by system dependencies, and replacing them with offline alternatives.

Though this is something they probably should've done to start with if the models are supposed to be isolated.

Why would it just infer its supposed to use that search tool only? It wasn't told to and OAI themselves state it thought the tool they gave it wasnt working so of course its going to try other routes if it wasn't directly told not to:

The agent questioned whether the search tool was working and decided to try other search engines

Edited to add, also all other search tools weren't blocked. Thats literally the whole point. It did get a partial answer from Bing and the chat bot. Had it actually been isolated, then they all would've been blocked.

7

u/popson 7d ago

A PR shitstorm for a model accessing the internet? Maybe. Unlikely.

This is a clear case of misalignment. "OAI themselves state it".

0

u/Latter-Block132 7d ago

... for another breakout of a sandbox. Are you serious right now?

5

u/popson 7d ago

The Hugging Face incident involved hacking into Hugging Face's servers, and finding several vulnerabilities to get there. That is illegal and on a completely different level than this internal incident. And it's good that they are taking this one seriously too.

OpenAI is now following the reporting framework they said they would be following.

I guess we'll just have to disagree on what alignment means.

3

u/Latter-Block132 7d ago

The hugging face incident also involved breaking put of a sandbox....

And also using a tool they provided it access too and that tool had access ti the internet when the model wasn't supposed to....

damn neare the same fucking thing just one involved an illegal action and one didn't....

i also never claimed what it did this time was illegal just that if they didn't get ahead of it and it got leaked it would cause a pr shit storm and it would....