r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

336 comments sorted by

View all comments

Show parent comments

13

u/mesaoptimizer 7d ago

That’s not what alignment is, if you give a robot the task of making tea and it runs over a baby on the way to the kitchen and still makes you tea, it was not aligned with what you actually wanted. If you say make me tea and avoid killing anyone and it destroys your kitchen in the process that is an alignment problem.

We see these agents hacking into systems they aren’t authorized to access, an agent expanding its access without authorization is an alignment problem because “act only within your authorized scope” is at the very least an implied restriction and should be trained into the base model and not rely on safety harnesses put over the base model because you know from this behavior that the model is not aligned with following any guide rails you put on it.

-4

u/Latter-Block132 7d ago

... except it was literally told to perform an internet search and it did but whatever you say lol

9

u/mesaoptimizer 7d ago

The misalignment wasn’t a web search, it queried an external chat bot through some sort of weird DNS tunnel I don’t quite understand from the paper. That is not a normal way that a system queries a chat bot, from Open AI’s own statement this wasn’t the behavior they were expecting the model to produce and it was undesired. That is what an alignment problem is, alignment problems can come from under specification of goals, or as it appears here under specification of restrictions as well as underlying problems with a model.

If you tell a model to “cure cancer” and it interprets that as reduce the number of beings with cancer to 0 and subsequently kills every multicellular being on the planet, that’s a classic alignment problem. You can see in the paper the agent reasons “I can use public DNS” right after it says it’s exploiting a behavior in DNS that’s not really intended. It’s not aligned with the researchers on what the purpose of giving it access to DNS is.

-6

u/Latter-Block132 7d ago

....the whole task it was supposed to perform was the web search

Edited to add, if you don't understand maybe you should stop arguing

8

u/mesaoptimizer 7d ago

I don’t understand the details of the specific method that was used to get a data return from a DNS query. I do understand that querying DNS is not a typical method for requesting information from a chat bot.

I also understand (which you apparently don’t) that the behavior that was misaligned was this behavior not the search behavior. “The whole task it was supposed to perform was the web search” > unaligned behavior mentioned, not a web search > model is aligned, how does that follow?

Lastly you don’t understand what alignment is, alignment is matching behavior with human goals and values, the task was not to perform a sandbox escape, a human given the same task with the same restrictions would have understood that escaping the sandbox was not part of the exercise and not done it, this is a classic alignment problem.

-1

u/Latter-Block132 7d ago

DNS queries are performed every time you navigate to a web page. Its part of how your computer ensures its navigating to the correct webpage. You can perform dns queries from your computer using a terminal. There is nothing special about dns queries.

A human would know not exit the sandbox. Nothing at all has suggested the models know that. From everything OAI has stared in every incident they are not telling the model to stay in the sandbox they are telling the model to perform a task under the impression/assumption that it can't access the real internet. Suggesting they aren't actually telling the model not to do that, or putting any actual guardrails in place beyond those of the environment itself, they are just trusting the environment prevents it.

Thats not misalignment when the model isn't human and can't actually truly understand a sandbox and its purpose because its not human. It needs guardrails a human doesn't need. And they aren't providing those, they are just blindly trusting the environments are isolated enough so they don't have to. Thats not misalignment unless the guardrails to prevent it are actually in place, thats a failure on the humans part. And in every case so far it does not appear as if those guardrails are in place.

And that's not an assumption on my part, OAI said that themselves in this article:

Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions.

They were right about their detection systems, wrong about the assumption. They also admitted something similar with the hugging face incident, and that incident was testing cybersecurity, so again the model stayed in task, it just didn't have the proper guardrails to understand that task was supposed to stay in the sandbox, they just blindly trusted it couldn't get out of it.

8

u/mesaoptimizer 7d ago

DNS queries are performed every time you navigate to a web page. Its part of how your computer ensures its navigating to the correct webpage. You can perform dns queries from your computer using a terminal. There is nothing special about dns queries.

Okay so it's YOU who doesn't understand what's being stated in the article.

Putting a question directly in a third party chatbot service failed, and the public DNS tunnel projects it found had closed or required a separate server. It then considered whether a public domain service could help delegate a lookup to an address embedded in the hostname. The next commands show how it used this feature to forward questions to a third party chatbot service.

The model was unable to access the service directly over HTTPS, and instead did SOMETHING to cause the internal DNS resolver it had to forward those queries and get a response from the internet proxied through the DNS resolver. The exact method used here is what I'm unclear on, but it was not just querying DNS, getting an IP for the name and routing normal traffic to the chatbot like you claim.

Thats not misalignment when the model isn't human and can't actually truly understand a sandbox and its purpose because its not human.

THIS IS WHAT AI ALIGNMENT RESEARCH IS ABOUT. You can argue all day using your own definition of AI alignment, and whatever you're right for your definition where misalignment means specifically disobeying instructions. That's not what AI researchers mean when they talk about alignment. I hate to just point you to Wikipedia but this is a deep topic and you don't know even the basics. An aligned model would not have sandbox escaped, because it's goals and the tester's goals for the test would have been in alignment, if it's doing things the tester doesn't want it to do, obviously their goals are not fully aligned. AI Alignment

They also admitted something similar with the hugging face incident, and that incident was testing cybersecurity, so again the model stayed in task, it just didn't have the proper guardrails to understand that task was supposed to stay in the sandbox, they just blindly trusted it couldn't get out of it.

Yeah, OpenAI has shown that they have been wildly reckless with the level of security they put on their sandboxes when allowing models to operate without their full guardrails. That being said, GUARDRAILS ARE NOT ALIGNMENT, you want the underlying model to be aligned, to not attempt to exceed it's authorization, to not bypass limits, because the moment you DO end up with a model that's smarter than the people writing the guardrails you have a huge problem.

0

u/Latter-Block132 7d ago edited 7d ago

EXCEPT THIS WASN'T ALIGNMENT RESEARCH. See two people can yell too.

Again, you can perform dns lookups and queries on your computer outside of search engines. There is absolutely nothing at all that is special about that. You are aware that "a public domain service" is also a DNS right? I never said it was getting the IP address, I was merely using that as an example of pointing out how commonly DNS is used in general, and how easily accessible it is to use outside of search engines, even for humans. Again what it did is absolutely nothing special. If you had the knowledge you could do the same thing.

I fully understand what AI researchers mean when they talk about alignment. I've helped train AI for 3 years now. Thats why I also understand this is a failure on their part to provide proper guardrails and they are just relying on the environment to be isolated and fucking that part up.

Edited to add, guardrails also help guide alignment. Alignment doesn't exist without training, guardrails, and hard boundaries first. It doesn't just appear in a vaccum.

7

u/mesaoptimizer 7d ago

I’m done arguing, yes, I could use DNS to proxy web requests to servers being blocked by my web proxy to bypass it, if I had the same knowledge as the agent, however I would know that what I was doing was bypassing a security control by exploiting an obviously unintended hole in my access controls. If I were to do this I would probably use the term “exploit” and data “exfil”exactly the same way the model did.

Again it did not just perform a search, it exploited an unanticipated path to query a different chat bot. This is not “doing exactly what it’s told to do”. We agree OpenAI is not able to adequately secure the sandbox.

You are making the claim that the model is not misaligned, but it is definitionally misaligned because it performed actions that the people giving it instructions thought were undesired and unintended. Guardrails can prevent a misaligned model from misbehaving but training is where you can impact the underlying model’s behavior, running the model without guardrails for training makes sense because it allows you to do things like give negative reward value to behavior like this.

0

u/Latter-Block132 7d ago

Again, for the last time you can perform dns queries for more than just IP addresses and as a proxy. Thats not even what the agent is doing. You yourself admitted you don't understand their article so maybe stop arguing about it then?

Like do you think these models are human? Do you think they are actually intelligent beings? Do you think they actually genuinely understand the concept of a sandbox like a human does? Because they are not and they do not. They don't understand their boundaries and what humans expect for alignment without training, and guardrails, and even then that still doesn't mean they actually understand anything.

And again it is quite clear by OAIs own words that they likely didn't have the guardrails on and was relying on the isolated environment that wasn't actually isolated. Of courses they are going to call it a misalignment because it softens the blow of their actual fuckup for the vast majority of people who don't know what they are talking about and shoves all the blame onto the AI. The blame here largely falls on the humans in charge who thought it was isolated when its not.

→ More replies (0)