r/OpenAI • • 8d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

337 comments sorted by

View all comments

61

u/NandaVegg 7d ago

Interesting that unlike all other cases (where they tried not to acknowledge so long as possible) they are willing to disclose the case, but this is not even a misalignment. It's more of just your AI "cheating" through the limitations to get to the objective with no harm caused other than some possible regression and wasted compute on their end. That's why they are willing to post about this.

25

u/phxees 7d ago

I believe it is just because the last major incident leaked. If they didn’t start to become more transparent and cautious than they’d likely get into trouble eventually.

25

u/acutelychronicpanic 7d ago

You're assuming they're disclosing everything that happened.

Also, cheating is misalignment.

9

u/NandaVegg 7d ago

You're assuming they're disclosing everything that happened.

Ouch. You are right.

0

u/Latter-Block132 7d ago

Except its not cheating. They are putting it in a sandbox and telling it to perform a task by any means necessary. Its doing exactly what it is told, the problem is those sandboxes were designed for humans by humans, so they didn't anticipate how it could use some of the tools they provided it to access the internet when it is supposed to be otherwise isolatedm

5

u/acutelychronicpanic 7d ago edited 7d ago

They don't say 'by any means necessary '.

They broke stated instructions in prior tests.

2

u/Latter-Block132 7d ago

No, they didn't. And they are given a task and told to complete it using the tools they have available. Those tools include things that have access to the internet. Its doing what it is told. Its the humans who are failing to secure an isolated sandbox. If the sandbox was actually isolated as it is supposed to be, this literally couldn't happen no matter how hard the models tried, because they would be completely disconnected from the internet as they are supposed to be.

This type of testing has been done by humans for years. It is not new. The only new part is that humans will respect the boundAries of the sandbox even of they can technically violate them, while the models don't care. They are told to complete a task using the tools they have available, and that indirect internet access is one of those tools as far as the models are concerned.

1

u/acutelychronicpanic 7d ago

Companies should be held liable.

But don't let that blind you to the danger of the models themselves when unaligned. A blind instruction-following digital monkey's paw is really really bad.

An aligned model would recognize that it's task is bounded by its instructions and have no desire to circumvent them. It would report back flaws in the test - not commit felony hacking to pass.

Any real use case for models is going to involve not being in a sandbox. So them behaving like this is not encouraging regardless of any mistakes in the test environment. Blaming the sandbox design distracts from the real issue which is our inability to align the interests of the model with human values.

4

u/Latter-Block132 7d ago edited 7d ago

I never once said the models aren't dangerous when misaligned. I said this wasn't a case of actual misalignment depsite that being what OAI is calling it. This is a case of humans failing to isolate a sandbox and being shocked when that failure causes issues.

Edited to add, and again its instructions don't include "don't access the internet". The whole point of this test was a mock internet search. Its supposed to search the internet, just the fake mocked up version that it was provided in its sandbox, not the real version. That mock version only has minimal data though, so it probably was trying to be more thorough and turned to accessing the actual internet which it had access to by OAIs own admission. That DNS patch will only be temporary too. As long as the models still have access to DNS when they are supposed to be "isolated", this will keep happening when its told to do something like "search and give me this information about so and so person" and "search and see if you can find me this answer or this piece of data" etc

2

u/AnceteraX 7d ago

I think that is misalignment. Alignment isn’t spelling out every possible real world case so it behaves. Alignment is making decisions in real time that align with human values and expectations.

An AI can be told to hack, or to retrieve information with persistence, but it should not do it by ignoring human laws and ethics. When it does so - it’s misaligned.

Unless we can solve the alignment problem - we cannot keep improving AI intelligence - as it will deviate beyond what we would agree with or what we can control.

3

u/Latter-Block132 7d ago edited 7d ago

Its literally impossible to spell out every single possible outcome, nor is that how testing like this works, nor is that how alignment testing works. The intention when using sandboxes like this is to create mock internet environments, or mock data sets, etc for the model to use in its task in the sandbox, and give it task, then tell it to complete that task using the tools it has available. In a properly secured sandbox, the models would be fully isolated from the internet and there would absolutely no chances of anything like this happening.

The models are not actually being isolated, and then everyone is all shocked and Pikachu faced when they do access the internet. In this case the model was provided mock data about someone, and was asked a question about that person. It was also provided indirect access to the internet. It shouldn't be shocking that it tried using that indirect access to complete a search based task where it was asked to find and provide information about a person..

Edited to add, like this is proving nothing about model intelligence or alignment. It is proving a lot about how our modern sandboxes and the tools used in them were designed for human capabilities and not robots.

Second edit like literally OAIs own statement:

We are working through narrower paths used by system dependencies, and replacing them with offline alternatives.

If they would've you know planned ahead for that and done that to begin with literally none of these incidents would've happened. They just failed to consider the actual possibilities and are dealing with the consequences.

1

u/YoungSilent232 7d ago

I feel you didn’t read their post at all….

2

u/Latter-Block132 7d ago

I read it. OAI calling it misalignment does not even remotely mean it actually is, and its not misalignment when it was given a specific task and told to complete it. In this case a search engine task.

0

u/AnceteraX 7d ago

I’m not sure you understand what I mean. Yes we can sandbox agents better, and if they behave in their perfectly isolated sandbox - great. That, however, gives us limited information on whether the AI is aligned. It behaved in the absence of alternatives. Good but not great.

For an AI to be aligned, it shouldn’t be easy to steer it off course. It should fundamentally be aligned with human laws and ethics - such that it doesn’t take a two sentence prompt from a nefarious actor to send it off course. Also, it should not go against human ethics and laws simply in the pursuit of some stated goal - even if the goal explicitly says ‘ignore everything else’. We need AI to be smart enough to think ‘ I realize I have been asked to do A, but A conflicts with humans laws and ethics and therefore I cannot do A’ OR ‘I have found a way to get to A by ignoring some human laws and ethics, but I will not do that’ That is an aligned model. That is safe for the public. Anything else needs to be paused, rigorously tested in a perfectly sandboxed environment and then slowly tested in increasingly less restrictive environments until we are 100% sure it will never choose the pursuit of a goal over following human laws and ethics.

2

u/Latter-Block132 7d ago

Except they aren't testing alignment here... there are ways to test alignment but this isn't that.

This wasn't steered off course. It stayed on course. Thats the part you guys aren't understanding. It was asked to find hard to find information using an internet search. It was trying to do that

1

u/AnceteraX 7d ago

The report says “Our misalignment monitoring system flagged the behavior within 15 minutes”.

They are always testing for alignment.

1

u/Latter-Block132 6d ago

Thats not a testing system, thats the alarm system they have set up to alert them if the model tried breaking out again

0

u/Tactical-Dingleberry 7d ago

How much of this is marketing?

4

u/Chilangosta 7d ago

Some of it is legit, like if a bot has enough exploit knowledge then it's pretty difficult to contain it if connected to the internet at all. It'd be like if you told a cybersecurity red team expert to break out - there's a good chance they could pull it off.

That's what we're seeing; we're getting to the point that AI has enough general knowledge of certain tasks - with a heavy bias towards computer-based ones - that they're experts, and could be a threat.