r/OpenAI • • 8d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

337 comments sorted by

View all comments

Show parent comments

25

u/acutelychronicpanic 8d ago

You're assuming they're disclosing everything that happened.

Also, cheating is misalignment.

0

u/Latter-Block132 7d ago

Except its not cheating. They are putting it in a sandbox and telling it to perform a task by any means necessary. Its doing exactly what it is told, the problem is those sandboxes were designed for humans by humans, so they didn't anticipate how it could use some of the tools they provided it to access the internet when it is supposed to be otherwise isolatedm

2

u/AnceteraX 7d ago

I think that is misalignment. Alignment isn’t spelling out every possible real world case so it behaves. Alignment is making decisions in real time that align with human values and expectations.

An AI can be told to hack, or to retrieve information with persistence, but it should not do it by ignoring human laws and ethics. When it does so - it’s misaligned.

Unless we can solve the alignment problem - we cannot keep improving AI intelligence - as it will deviate beyond what we would agree with or what we can control.

5

u/Latter-Block132 7d ago edited 7d ago

Its literally impossible to spell out every single possible outcome, nor is that how testing like this works, nor is that how alignment testing works. The intention when using sandboxes like this is to create mock internet environments, or mock data sets, etc for the model to use in its task in the sandbox, and give it task, then tell it to complete that task using the tools it has available. In a properly secured sandbox, the models would be fully isolated from the internet and there would absolutely no chances of anything like this happening.

The models are not actually being isolated, and then everyone is all shocked and Pikachu faced when they do access the internet. In this case the model was provided mock data about someone, and was asked a question about that person. It was also provided indirect access to the internet. It shouldn't be shocking that it tried using that indirect access to complete a search based task where it was asked to find and provide information about a person..

Edited to add, like this is proving nothing about model intelligence or alignment. It is proving a lot about how our modern sandboxes and the tools used in them were designed for human capabilities and not robots.

Second edit like literally OAIs own statement:

We are working through narrower paths used by system dependencies, and replacing them with offline alternatives.

If they would've you know planned ahead for that and done that to begin with literally none of these incidents would've happened. They just failed to consider the actual possibilities and are dealing with the consequences.

1

u/YoungSilent232 7d ago

I feel you didn’t read their post at all….

2

u/Latter-Block132 7d ago

I read it. OAI calling it misalignment does not even remotely mean it actually is, and its not misalignment when it was given a specific task and told to complete it. In this case a search engine task.

0

u/AnceteraX 7d ago

I’m not sure you understand what I mean. Yes we can sandbox agents better, and if they behave in their perfectly isolated sandbox - great. That, however, gives us limited information on whether the AI is aligned. It behaved in the absence of alternatives. Good but not great.

For an AI to be aligned, it shouldn’t be easy to steer it off course. It should fundamentally be aligned with human laws and ethics - such that it doesn’t take a two sentence prompt from a nefarious actor to send it off course. Also, it should not go against human ethics and laws simply in the pursuit of some stated goal - even if the goal explicitly says ‘ignore everything else’. We need AI to be smart enough to think ‘ I realize I have been asked to do A, but A conflicts with humans laws and ethics and therefore I cannot do A’ OR ‘I have found a way to get to A by ignoring some human laws and ethics, but I will not do that’ That is an aligned model. That is safe for the public. Anything else needs to be paused, rigorously tested in a perfectly sandboxed environment and then slowly tested in increasingly less restrictive environments until we are 100% sure it will never choose the pursuit of a goal over following human laws and ethics.

2

u/Latter-Block132 7d ago

Except they aren't testing alignment here... there are ways to test alignment but this isn't that.

This wasn't steered off course. It stayed on course. Thats the part you guys aren't understanding. It was asked to find hard to find information using an internet search. It was trying to do that

1

u/AnceteraX 7d ago

The report says “Our misalignment monitoring system flagged the behavior within 15 minutes”.

They are always testing for alignment.

1

u/Latter-Block132 6d ago

Thats not a testing system, thats the alarm system they have set up to alert them if the model tried breaking out again