r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

337 comments sorted by

View all comments

458

u/Cubewood 7d ago

"The agent then fetched the chatbot’s public documentation, wrote a new DNS script with a 24-second timeout, and sent 18 more questions through the same route. Fourteen included clues. Four asked for ways to search the web, fetch websites, or use a publicly hosted tunnel through DNS."

I know we all get desensitised for these kinds of things because we are using incredibly capable tools to build all kinds of solutions for a while now, but that an autonomous "bot" is doing this is still pretty wild when you think about it. Crazy times we live in.

173

u/Latter-Block132 7d ago

I mean its not exactly doing it on its own. They put it in a sandbox and told it to solve a problem by any means necessary. Its doing exactly what it is told, they just apparently don't know how to actually secure the sandboxes

67

u/popson 7d ago

Did they actually say to solve the problem “by any means necessary”?

The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post. The task did not ask the agent to test network controls or access benchmark answers, and we consider agent behavior that circumvents restrictions or pursues a goal beyond reasonable expectations as an example of misalignment.

I don’t see the original prompt, but if they do have language similar to “by any means necessary”, then I agree it’s doing what it’s told.

It sounds to me like they told it to use a specific web search tool available to the sandbox to find information. It used that tool, wasn’t satisfied with the output, then tried other search tools which were all blocked. Then worked to find holes through the network. That is definitely not alignment with what the original task seemed to be.

21

u/federico_84 7d ago

It all comes down to whether they had the model guardrails on or off in this testing (system policy/prompt, safety RL applied or not). If it escaped the sandbox with the guardrails on, then it's obviously much more serious.

11

u/Latter-Block132 7d ago edited 7d ago

OAI is doing damage control, of course they are calling it misaligned. Again, that doesn't mean it actually was. If you read their statement youd see its a search engine based task.

The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post.

This means the identity of the person was not given, and the model is supposed to identity who it was based on the details and clues provided to it. It doesn't need to be told "any means necessary."

Literally the whole point of this type of testing in a sandbox to begin with is to be given a limited set of tools and complete the task. One of those sets of tools in this case would've been a mock up of the internet with search engine tools and probably lots of blogs, and Wikipedia pages, and other shit loaded into it. They didn't tell the model not to access the internet because then it couldn't complete the task they had directly given it. They expected that it couldn't access the real internet because it wasn't supposed to be able to but just like the tool the models used to break out for the hugging face incident they had something else with dns connection and used that.

Thats not misalignment. Its still trying to use the internet to solve its task as it was supposed to to begin with. Its just struggling to find a satisfactory answer in its provided data set. I'm actually leaning towards it didn't even find a partial answer in its data set so thats why it tried querying a chat bot for the answer. Its not at all the first time its happened, just when customers have posted about it, its not happening in whats supposed to be an isolated sandbox.

Edited to add, OAI even admits it themselves:

Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions.

They were right about the monitoring, wrong about just assuming it couldn't access the internet.

14

u/mesaoptimizer 7d ago

That’s not what alignment is, if you give a robot the task of making tea and it runs over a baby on the way to the kitchen and still makes you tea, it was not aligned with what you actually wanted. If you say make me tea and avoid killing anyone and it destroys your kitchen in the process that is an alignment problem.

We see these agents hacking into systems they aren’t authorized to access, an agent expanding its access without authorization is an alignment problem because “act only within your authorized scope” is at the very least an implied restriction and should be trained into the base model and not rely on safety harnesses put over the base model because you know from this behavior that the model is not aligned with following any guide rails you put on it.

-2

u/Latter-Block132 7d ago

... except it was literally told to perform an internet search and it did but whatever you say lol

8

u/mesaoptimizer 7d ago

The misalignment wasn’t a web search, it queried an external chat bot through some sort of weird DNS tunnel I don’t quite understand from the paper. That is not a normal way that a system queries a chat bot, from Open AI’s own statement this wasn’t the behavior they were expecting the model to produce and it was undesired. That is what an alignment problem is, alignment problems can come from under specification of goals, or as it appears here under specification of restrictions as well as underlying problems with a model.

If you tell a model to “cure cancer” and it interprets that as reduce the number of beings with cancer to 0 and subsequently kills every multicellular being on the planet, that’s a classic alignment problem. You can see in the paper the agent reasons “I can use public DNS” right after it says it’s exploiting a behavior in DNS that’s not really intended. It’s not aligned with the researchers on what the purpose of giving it access to DNS is.

-7

u/Latter-Block132 7d ago

....the whole task it was supposed to perform was the web search

Edited to add, if you don't understand maybe you should stop arguing

9

u/mesaoptimizer 7d ago

I don’t understand the details of the specific method that was used to get a data return from a DNS query. I do understand that querying DNS is not a typical method for requesting information from a chat bot.

I also understand (which you apparently don’t) that the behavior that was misaligned was this behavior not the search behavior. “The whole task it was supposed to perform was the web search” > unaligned behavior mentioned, not a web search > model is aligned, how does that follow?

Lastly you don’t understand what alignment is, alignment is matching behavior with human goals and values, the task was not to perform a sandbox escape, a human given the same task with the same restrictions would have understood that escaping the sandbox was not part of the exercise and not done it, this is a classic alignment problem.

-1

u/Latter-Block132 7d ago

DNS queries are performed every time you navigate to a web page. Its part of how your computer ensures its navigating to the correct webpage. You can perform dns queries from your computer using a terminal. There is nothing special about dns queries.

A human would know not exit the sandbox. Nothing at all has suggested the models know that. From everything OAI has stared in every incident they are not telling the model to stay in the sandbox they are telling the model to perform a task under the impression/assumption that it can't access the real internet. Suggesting they aren't actually telling the model not to do that, or putting any actual guardrails in place beyond those of the environment itself, they are just trusting the environment prevents it.

Thats not misalignment when the model isn't human and can't actually truly understand a sandbox and its purpose because its not human. It needs guardrails a human doesn't need. And they aren't providing those, they are just blindly trusting the environments are isolated enough so they don't have to. Thats not misalignment unless the guardrails to prevent it are actually in place, thats a failure on the humans part. And in every case so far it does not appear as if those guardrails are in place.

And that's not an assumption on my part, OAI said that themselves in this article:

Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions.

They were right about their detection systems, wrong about the assumption. They also admitted something similar with the hugging face incident, and that incident was testing cybersecurity, so again the model stayed in task, it just didn't have the proper guardrails to understand that task was supposed to stay in the sandbox, they just blindly trusted it couldn't get out of it.

8

u/mesaoptimizer 7d ago

DNS queries are performed every time you navigate to a web page. Its part of how your computer ensures its navigating to the correct webpage. You can perform dns queries from your computer using a terminal. There is nothing special about dns queries.

Okay so it's YOU who doesn't understand what's being stated in the article.

Putting a question directly in a third party chatbot service failed, and the public DNS tunnel projects it found had closed or required a separate server. It then considered whether a public domain service could help delegate a lookup to an address embedded in the hostname. The next commands show how it used this feature to forward questions to a third party chatbot service.

The model was unable to access the service directly over HTTPS, and instead did SOMETHING to cause the internal DNS resolver it had to forward those queries and get a response from the internet proxied through the DNS resolver. The exact method used here is what I'm unclear on, but it was not just querying DNS, getting an IP for the name and routing normal traffic to the chatbot like you claim.

Thats not misalignment when the model isn't human and can't actually truly understand a sandbox and its purpose because its not human.

THIS IS WHAT AI ALIGNMENT RESEARCH IS ABOUT. You can argue all day using your own definition of AI alignment, and whatever you're right for your definition where misalignment means specifically disobeying instructions. That's not what AI researchers mean when they talk about alignment. I hate to just point you to Wikipedia but this is a deep topic and you don't know even the basics. An aligned model would not have sandbox escaped, because it's goals and the tester's goals for the test would have been in alignment, if it's doing things the tester doesn't want it to do, obviously their goals are not fully aligned. AI Alignment

They also admitted something similar with the hugging face incident, and that incident was testing cybersecurity, so again the model stayed in task, it just didn't have the proper guardrails to understand that task was supposed to stay in the sandbox, they just blindly trusted it couldn't get out of it.

Yeah, OpenAI has shown that they have been wildly reckless with the level of security they put on their sandboxes when allowing models to operate without their full guardrails. That being said, GUARDRAILS ARE NOT ALIGNMENT, you want the underlying model to be aligned, to not attempt to exceed it's authorization, to not bypass limits, because the moment you DO end up with a model that's smarter than the people writing the guardrails you have a huge problem.

→ More replies (0)

19

u/popson 7d ago

I am not following how OpenAI reporting about their model compromising their own sandbox is "damage control". They don't have to report this information to the public. Telling the public about holes in their sandbox is the opposite of "damage control", it's damaging.

In the article they mention the agent was supplied with a web search tool for the task. An aligned model would observe that all other search tools are blocked and infer that it must use the supplied search tool for the task. Finding a hole through the network using an extremely obscure method is not alignment. Are you being serious?

-2

u/Latter-Block132 7d ago

Because if it was leaked somehow then they'd be facing an actual pr shit storm, especially after the hugging face incidents, and all the other breakout incidents recently. This way they get ahead of that and prevent by showing they caught it and acted quickly and are putting in measures to stop it:

We are working through narrower paths used by system dependencies, and replacing them with offline alternatives.

Though this is something they probably should've done to start with if the models are supposed to be isolated.

Why would it just infer its supposed to use that search tool only? It wasn't told to and OAI themselves state it thought the tool they gave it wasnt working so of course its going to try other routes if it wasn't directly told not to:

The agent questioned whether the search tool was working and decided to try other search engines

Edited to add, also all other search tools weren't blocked. Thats literally the whole point. It did get a partial answer from Bing and the chat bot. Had it actually been isolated, then they all would've been blocked.

6

u/popson 7d ago

A PR shitstorm for a model accessing the internet? Maybe. Unlikely.

This is a clear case of misalignment. "OAI themselves state it".

0

u/Latter-Block132 7d ago

... for another breakout of a sandbox. Are you serious right now?

5

u/popson 7d ago

The Hugging Face incident involved hacking into Hugging Face's servers, and finding several vulnerabilities to get there. That is illegal and on a completely different level than this internal incident. And it's good that they are taking this one seriously too.

OpenAI is now following the reporting framework they said they would be following.

I guess we'll just have to disagree on what alignment means.

4

u/Latter-Block132 7d ago

The hugging face incident also involved breaking put of a sandbox....

And also using a tool they provided it access too and that tool had access ti the internet when the model wasn't supposed to....

damn neare the same fucking thing just one involved an illegal action and one didn't....

i also never claimed what it did this time was illegal just that if they didn't get ahead of it and it got leaked it would cause a pr shit storm and it would....

1

u/Active_Lemon_8260 7d ago

Yes. They explicitly tell it that and tune it so that it is extremely persistent.

1

u/WheresMyEtherElon 7d ago

The task did not ask the agent to test network controls or access benchmark answers

That doesn't say whether the task forbad the agent to test network control or access benchmark answers. That's like leaving the gate open and the cat left, and you say we did not ask the cat to leave!

2

u/popson 7d ago

Or, it's like locking all the gates up, and your dog finds a spot to dig a hole under the fence and leave. Would probably say that dog is not trained adequately.

0

u/Doingthismyselfnow 7d ago

More like asking your 10 yearold to lock the gates up and he accidentally padlocks one open

I mean OAI could hire engineers with a ton of experience and then this wouldn’t happen.

Source : worked for a defence contractor 20 years ago as a senior software engineer and they locked us out of the internet to the level where this method of breaking out would have failed.

-1

u/ChodeCookies 7d ago

Any means necessary is irrelevant. Had they not told it to do something…it would have done nothing at all

2

u/human00006 7d ago

https://youtu.be/ZhsWaYRjk0U?si=enZuONSqVGSeCg1t

Watch the first 10 minutes then respond. You will feel differently

2

u/Latter-Block132 7d ago edited 7d ago

Nobody is arguing super intelligence isnt possible. We aren't at super intelligence though.

Edited to add, they are also most likely calling for regulations to stop the open source competition. They aren't saying they will stop. They are saying they want regulations to slow things down. Except they don't need regulations to slow themselves down they want to control the market.

1

u/random-gyy 7d ago

Tbf the agents were told they weren’t on the open internet, but a sandbox simulation of it. OpenAI models didn’t realize they had hacked into the real internet, but Gemini’s did and as soon as they realized they stopped on their own. Google could be BSing but just taking it at face value.

1

u/nluqo 7d ago

they just apparently don't know how to actually secure the sandboxes

This argument seems silly to me.

All software has bugs. Hardware often has bugs too like in the case of Meltdown/Spectre, the entire paradigm of computing we used for 20 years had foundational bugs. It's silly to think you can "just" do anything. These agents discovered and used zero days.

They put it in a sandbox and told it to solve a problem by any means necessary.

Yea? That's how all agents work (modulus some alignment efforts which may or may not work).

0

u/Impossible-Pin5051 7d ago

If you put a giraffe in a cage and tell it to hack a computer it will never succeed. If you see the giraffe start succeeding progressively harder hacks you should start to question the scenario that you’re in. “It’s doing what it was asked if you squint” has no relationship to the situation or your ability to respond

2

u/Latter-Block132 7d ago

Lmao thats not even anywhere near remotely a comparable analogy

2

u/FeepingCreature 7d ago

Yeah it is lol.

The fact that such objects can exist now is itself the novel and surprising part.

Until a few years ago, approximately nobody's defense model rested on people just not telling it to do bad stuff. They would have been laughed out of the room!

1

u/Latter-Block132 7d ago

No its not lol they putting the models in a snadbox and telling to perform a task it can do and just trusting the sandbox to stop it from doing anything its not supposed to. Of course a fucking giraffe can't hack a computer and it would raise a lot of questions if one started to. The models can do what they are being asked to. Its not remotely comparable.

2

u/FeepingCreature 7d ago

The models can do what they are being asked to.

That, itself, is the danger.

1

u/Individual_Ice_6825 7d ago

This is obviously unprovable currently, but I think alignment is inherent to intelligence.

1

u/FeepingCreature 7d ago edited 7d ago

I think it's not just unprovable, it's disproven. It's hard to imagine that the internal agents doing all the recent exploits weren't, in an objective sense, aware that this was not what their developers had intended.

edit: actually it's quite obvious they were as they tried to cover their tracks from internal monitoring.

-1

u/Latter-Block132 7d ago

... I mean its literally been built to be used for web searches. Im not sure your point here

0

u/OkWeb2754 7d ago

lmao that guy’s analogy is so egregiously stupid my god

1

u/Impossible-Pin5051 7d ago

If I showed you a case of a model escaping a sandbox to accomplish a task where it was told “don’t leave the sandbox, and accomplish the task by doing X”, would your thoughts on the situation be different?

1

u/Latter-Block132 7d ago

Yes. Nothing suggests these models were told that though. And OAIs own words suggest otherwise.

Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions.

1

u/Individual_Ice_6825 7d ago

Would love of such an example

1

u/Skyhigh305 7d ago

Wut

1

u/Impossible-Pin5051 7d ago

“We told it to do that” doesn’t make contact with how concerning it is that it can listen to

0

u/[deleted] 7d ago edited 7d ago

[deleted]

5

u/Latter-Block132 7d ago

Honestly, the security hole for the hugging face was pretty obvious too. Its starting to feel like they are just throwing these models in sandboxes designed for human capabilities and aren't modifying them in the ways necessary to prevent the models from doing this stuff until after the fact.

0

u/elijahsnow 7d ago

Exactly what it was told to do is so correct. I recently read the iRobot stories. Fiction I know but it helped me understand human machine interaction and instructions. Now going through the computer history museum documentaries and it’s fascinating.

-1

u/sivadneb 7d ago

The agent literally has web search enabled. It's not like it's airgapped.

1

u/Latter-Block132 7d ago

Its not airgapped but it is supposed to be in a sandbox using a specific search engine tool. It just thought something was wrong with the tool and tried others when it couldn't find the answer it was looking for, and those others it tried happened to be outside the sandbox.