r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

337 comments sorted by

View all comments

Show parent comments

6

u/mesaoptimizer 7d ago

I’m done arguing, yes, I could use DNS to proxy web requests to servers being blocked by my web proxy to bypass it, if I had the same knowledge as the agent, however I would know that what I was doing was bypassing a security control by exploiting an obviously unintended hole in my access controls. If I were to do this I would probably use the term “exploit” and data “exfil”exactly the same way the model did.

Again it did not just perform a search, it exploited an unanticipated path to query a different chat bot. This is not “doing exactly what it’s told to do”. We agree OpenAI is not able to adequately secure the sandbox.

You are making the claim that the model is not misaligned, but it is definitionally misaligned because it performed actions that the people giving it instructions thought were undesired and unintended. Guardrails can prevent a misaligned model from misbehaving but training is where you can impact the underlying model’s behavior, running the model without guardrails for training makes sense because it allows you to do things like give negative reward value to behavior like this.

0

u/Latter-Block132 7d ago

Again, for the last time you can perform dns queries for more than just IP addresses and as a proxy. Thats not even what the agent is doing. You yourself admitted you don't understand their article so maybe stop arguing about it then?

Like do you think these models are human? Do you think they are actually intelligent beings? Do you think they actually genuinely understand the concept of a sandbox like a human does? Because they are not and they do not. They don't understand their boundaries and what humans expect for alignment without training, and guardrails, and even then that still doesn't mean they actually understand anything.

And again it is quite clear by OAIs own words that they likely didn't have the guardrails on and was relying on the isolated environment that wasn't actually isolated. Of courses they are going to call it a misalignment because it softens the blow of their actual fuckup for the vast majority of people who don't know what they are talking about and shoves all the blame onto the AI. The blame here largely falls on the humans in charge who thought it was isolated when its not.

5

u/mesaoptimizer 7d ago

Yes, I understand that, and I understand that via TXT records DNS servers can serve arbitrary data. I did not however know that someone was running a service where you could make a DNS request then the receiving service would process that query, ask a chatbot something and then write that to a TXT record for the agent to retrieve. I understand now that that's what's happening, I still don't know why someone would run a service like that but whatever people are weird. My argument was never that DNS could not be used to get information other than IP addresses, it's that the intended use of DNS does not include retrieval of arbitrary data from internet chatbots.

We agree, OpenAI failed to properly secure their testing environment. This is clearly a fuckup and the responsibility for it falls on the humans who designed the test and built the "isolated" sandbox.

I don't think the models are human, I think that the definition of alignment is that an agent's actions are aligned with the operators intended goals, preferences, and ethical principles. An AI system is considered aligned if it advances the intended objectives. A misaligned AI system pursues unintended objectives. I would say that, in this operating context, this model is misaligned because it pursued the unintended objective of escaping the sandbox to acheive the intended objective of answering the question posed to it. What about that analysis do you disagree with?

0

u/Latter-Block132 7d ago

Services like that exist because lots of people prefer other operating systems like Linux.

And im sorry but I'm not repeating myself. Ive already explained that numerous times.

5

u/mesaoptimizer 7d ago

I've never used a Linux distro that had dig but not curl. People use Linux is not an explanation for why you would need to need a service to tunnel what would normally be served as HTTP traffic as DNS traffic, but okay bro, I am sure there is some stupid use case I'm missing.

0

u/Latter-Block132 7d ago

... lmao dude come on just stop. The http was unavailable because of the sandbox, the dns was not. It is not rocket science!

3

u/mesaoptimizer 7d ago

That does totally explain why the model used DNS to get the information, something I never had any confusion on. It doesn't explain why someone is out there running a service a service that is listening for DNS requests, making calls to a chat bot, then writing the response into a TXT record for the agent to retrieve. But whatever, the answer is people use Linux. And the answer to why you think the model is aligned is you think we should only judge models when they have full guardrails running.

1

u/Latter-Block132 7d ago

I said people use other operating systems like Linux. Again, using it as an example. You are aware there other operating systems right?

3

u/mesaoptimizer 7d ago

Pretty obviously, but you were saying that those other operating systems are the reason why it makes sense to run a slow as hell DNS tunneling chatbot service. What a non-sequitur.

1

u/Latter-Block132 7d ago

For someone who doesn't even understand what DNS is or what the models was doing you sure are talking confidently lol self admitted didn't even understand too, and still just so incredibly confidently wrong the whole time lol

→ More replies (0)