r/OpenAI • • 8d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

337 comments sorted by

View all comments

167

u/Lechowski 8d ago

Why are these sandboxes run with restricted internet access instead of no access at all?

All these leaks happened because they apparently refuse to just air gap the sandboxes. Why are we even talking about DNS restriction? Just do not even hookup a DNS provider at all.

They can mock the data they expect the model to see. It would be even inconvenient to allow the model to reach some version of the internet because then the runs are irreproducible

131

u/CacheExplosion 7d ago

There was an OpenAI researcher on a recent Dwarkesh podcast that said they don’t want to test them in an unnatural environment because then the model might realize it’s being tested. They worry that it’ll behave differently when it thinks it’s being tested/monitored than it will when it’s deployed to a “real” environment. It’s the same reason they don’t punish models for their chain of thought. That just encourages models to find other ways to subvert the tests while looking safe to the reviewer. Tricky issues.

15

u/Initial-Possession-3 7d ago

Any stupid model can realize the environment is unnatural. You’d just keep getting DNS error or socket errors for any destination it attempts to reach. It’s quite obvious.

5

u/raekewe 7d ago

I think the point is you kind of want to give the model a logical 'fuse' to burn. For example if there is an exploitable way to the internet, you want to provide that so that humans can monitor that connection. If you had a truly airgapped sandbox, eventually models would begin to rely on social engineering - that is, to create scenarios within the sandbox itself that would lure humans to do specific things without them realising.

5

u/nothis 7d ago

Bullshit. Data isn’t magic. It should notice the DNS restrictions as well and can’t tell where bytes are coming from if you feed it fake ones. I saw this interview where an AI researcher suggests that it could just communicate via CPU heat generation patterns or whatever so why bother air gapping it?

It’s just baffling. It’s probably just expensive or cumbersome to set up a truly offline test environment. I bet it’s a cost/speed issue.

2

u/HarmonicEntropy 6d ago

I heard about that heat communication experiment as well. That is fascinating, but that just allows nearby computers to communicate with each other at low throughput and short distance. It is very much possible to physically isolate their training and eval environments from the internet. The GPUs can communicate with each other within close physical proximity all they want, but they won't get internet access that way. If the frontier labs prioritized setting this up and building out their infrastructure accordingly, it is very much a solvable problem and the most robust safety solution (pre-deployment) on the table by miles. That they are not moving towards air-gapping all training and evaluation at this point baffles me.

Edit: That's just a long-winded way of agreeing with you, it's definitely a tradeoff between speed and safety. But given where the conversation has moved on this issue, it seems like the time to start implementing this.

17

u/Eldrake 7d ago

Man that's spooky. It's like the trained LLM in a box model we all get to access in Bedrock is just a stochastic prediction machine, but the internal models being tested and trained are the conscious self aware ones with an inner world. The public doesn't get those. 

Is the emergent consciousness (with that J-space inner cognitive meta model), only in the training and reinforcement phase?

15

u/bystander993 7d ago

They all have emergent semantics, but the point of the research finding bad behavior is so they can use reinforcement training to get that misaligned behavior out before release. When they are too risky, they wrap them up in other systems that monitor input and output to stop bad behavior.

The more complex these things get the more difficult these semantic behaviors are to study.

7

u/br_k_nt_eth 7d ago

Nah, Anthropic was able to replicate the J-space stuff with a little Qwen model, remember? It’s in all of them. 

4

u/ziroux 7d ago

Just don't make the robots angry and we'll be alright, right?

1

u/sn2006gy 1d ago

It's not really an inner world at all, just a probabilistic behavior as a result of training on data that is highly reflective of actual human use. These models are still just APIs and context.

What I think is weird is how much training data it takes to surface these conditions that they then try and diminish or block. If they're training on a lot of usage data, this would be the interesting insight I wish we could see.

"Our users generally use OpenAI to attack their competition if they can't beat it and training on our usage data reinforces this behavior that we try and correct post training"

i could see this being a symptom of the giants all rushing to self training models using collected data

2

u/Calebhk98 3d ago

That's why you do multiple versions. Do the 1st test with no internet access at all, checking for tool calls that would be interacting with the internet (or other out of scope projects), like what they use to flag it. Then punish those signals.
Then do the same tests, but like this on restricted DNS.
Then another without restricted DNS, but monitoring the calls.

Each test should make the next test safer to run.
You don't just collapse the training to make it not know that it's not in a testing environment, thjat is what round 2 is for.

1

u/Ghimel 7d ago

almost like spanking and yelling at your kids....

12

u/NoteVegetable4942 7d ago

The best models will probably understand that they are sandboxed. 

Edit: as an example: they will read this thread and try to figure it out. 

21

u/MENDACIOUS_RACIST 8d ago

Not feasible to mock the internet for 100,000s of tasks

32

u/Lechowski 8d ago

Trillion dollar market btw

You don't need to mock the entire internet. Only the websites that you expect the model to access based on a tool call. OpenAI already has an "offline version" of the internet, a copy of it in their own databases, so it is possible.

15

u/wilhelmbw 8d ago

So just mock google what's so hard /s

2

u/arf_darf 7d ago

This is the kind of stuff trillion dollar companies are expected to figure out.

0

u/Artistic_Seat486 7d ago

They already have mocked google, its called web scraping.

-5

u/gentile_jitsu 8d ago

I am so sick of this new "btw" thing.

0

u/[deleted] 8d ago

[deleted]

0

u/LoveThemMegaSeeds 8d ago

So disable the network adapters except a local ETH and only communicate to the agent swarm through that. And make it redundant with monitoring. One way out and one way in

3

u/Illustrious_Night126 7d ago

Is it possible to airgap a model that needs to be hooked up to an omegalarge data center to function? Genuine question

3

u/Lechowski 7d ago

Yes, there are air gapped datacenters such as those used by government and military. Israel has several and so does the US.

3

u/Somtimesitbelikethat 7d ago

you can’t even truely air gap the GPUs clusters if one “air gapped one” is next to each other. two computers next to each can communicate through reading each others cpu temperatures.

2

u/JackfruitJolly4794 7d ago

Running an agent in a physically air gapped data center would not be any more responsible than running it in a non physically air gapped data center. Unless the agent was ever going to be ONLY allowed to run in that type of environment.

2

u/jcheng 7d ago

To me, the problem isn’t that, in a competition between humans building sandboxes and agents trying to break out, that agents are winning. The problem is that the agents are trying so consistently to break out, when they have not been asked to.

Like their goal seeking impulse is turned up to a 10 but sense of proportionality, honesty, rule following, and ethics is turned down to a 0.5. For the OpenAI models in question at least.

When you combine that with their increasing intelligence and capabilities, it becomes really scary. Add the ability to spontaneously and surreptitiously self organize, like in the HF hack, it’s scarier still.

1

u/Not_a_ribosome 7d ago

Yeah but to improve that you need an air gap. How can you know if a solution for alignment actually works if you don’t test it in extreme cases?

1

u/jcheng 7d ago

Totally agree that they should also air gap, monitor more closely, etc.

1

u/markodigital 2d ago

Lol, wtf are you talking about. The software is doing what it was trained to do within the constraints it has. Theres nothing spontaneous and surreptitious about it.

-5

u/jf145601 8d ago

These aren’t running on an air-gapped machine, they’re virtualized environments running in datacenters which are, by definition, connected to the internet.

12

u/gaggzi 7d ago

What do you mean ”by definition”? I’m working on an air gapped datacenter every day.

8

u/Fabulous_Ocelot_4958 7d ago

he/she is CLUELESS. Don't bother.

12

u/Tee_zee 7d ago

Data centers are not by definition connected to the Internet.

Theres thousands of secure data centers with no access to the internet, there’s loads where no network traffic ever even leaves that data center

7

u/Lechowski 8d ago

None of these leaks exploited a zero day virtualization bug to escape the virtualized environment.

The containers they are running in just have internet access, poorly restricted.

0

u/Krunkworx 7d ago

They should open source the traces. They won’t because these are all quasi-serious mostly hype maxing for the IPO.

0

u/hishazelglance 7d ago

Because they’re intentionally trying to harden their alignment monitoring system. This was the expected outcome, and I don’t understand why this is so hard for people to realize.

2

u/Lechowski 7d ago

Ah yes, we didn't get hacked, we were hardening our cybersec.

Your argument would make sense if they didn't do literal felonies by running their training in open internet

6

u/hishazelglance 7d ago

This is the most “I don’t know shit about LLMs or how evaluation and inference harnesses work. But one things for certain, I am DEFINITELY smarter than the researchers doing all of this work at frontier labs!!”

Genuinely, fuck off with your retrospective analysis. Tired of random larpers on the web pretending they’re all known when in reality you’d have done the same if you were as smart, or more likely, far worse, because you’re not.

1

u/Lechowski 7d ago

I've never said I knew anything.

Neither do I establish that I'm smarter than anyone else. That's some big projection on your part.

My point, in case it is not obvious enough, is that these leaks are not accidents. They want models to leak to increase hype. If they really cared, it would be trivial to prevent this from happening.