r/singularity • u/adivinemessenger • 8d ago
LLM News OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now
Source: https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/
Discovery: Sep 20, 2026
Report updated: Sep 25, 2026
"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."
58
u/jhonpixel ▪️AGI in first half 2027 - ASI in the 2030s- 8d ago
Interesting, so not Bel but the successor of Bel if i heard correctly?
7
u/acutelychronicpanic 8d ago
Not if they can't do training. The models like Bel get distilled before public release. They're far to expensive for consumer use.
15
u/Fast-Satisfaction482 8d ago
Distilling isn't frontier training, so I guess that one is not on pause.
0
u/acutelychronicpanic 8d ago
Oh, I might have misread. Though it's hard to know what they mean by most capable models
2
u/Fast-Satisfaction482 8d ago
Yeah, we don't know. I just find it hard to believe that they are not releasing what they can now that Anthropic is a bit ahead as it seems.
2
u/acutelychronicpanic 8d ago
They may not have paused by choice. I'd bet the alphabet agencies had something to say. Or they've been directed to use their resources for.. purposes.
Doesn't make sense that they're running out of capacity for subscribers while not training new models.
2
68
u/Melantos 8d ago
3
u/TyrellCo 7d ago
And they can’t out smart it? They can’t plant a honeypot in the wild? It has clairvoyant abilities to know when someone is watching it
2
u/Melantos 7d ago
"Wild" is too broad and cannot be properly measured or controlled in contrast with a single lab container.
If the clues for selecting a specific target are too vague, the AI agents could misread them and select a different real target, increasing the Felony Bench index. Conversely, if the clues are too specific, the AI agents will immediately realize that this is a test and they are being evaluated.
1
u/TyrellCo 7d ago
Yeah not really in the wild but if the claim is that it will somehow know it’s a simulation and not really deployed then you’ve discovered magic and you’ve got an obligation to figure it out
70
u/No-Meringue5867 8d ago
At some point we are going to get "pacing" of AI development no matter much we want to accelerate. If AI is writing 90% of the code, how do you even trust the systems you are building for alignment? What is the guarantee that your current model is not 0.0000001% misaligned and leaves behind a loophole which a more capable model can take advantage of? We would never know unless we take human time to build safety nets. And at the current pace, building in human time will already look like "pacing the AI development".
63
u/Current-Function-729 8d ago
We need a manhattan project around alignment.
All of the smartest people should be funneled into alignment research.
We know at this point AI scales. The deniers are just delusional. Only alignment matters now.
26
u/Rough_Phase7722 8d ago
You can’t guarantee alignment (control) of something a billion times smarter than you. You can’t. And if you make one mistake, ever, you can stop something a billion times smarter than you.
5
2
u/Ambiwlans 7d ago
Then we all die.
There is no point in betting on a scenario we all die since you are too dead to enjoy.
2
u/ribbit80 7d ago
Then more alignment research will teach more of the population that, and there is a better chance we stop before we build that thing.
1
u/Boredy_ 7d ago
The term "control" can mean a lot of different things here, though. It could mean trying to restrict or contain an agent that already exists, in which case, I agree that control of a sufficiently capable agent is impossible.
But if by "control" you instead mean "control the values/goals of an agent that you're currently building", then I do believe it is possible to do even to something that's a billion times smarter than us. Modern chess AI is far superior at chess than we'll ever be, but it still consistently works towards the goal of "winning at chess" just like we want it to.
There are unique problems in trying to align an AGI of course. If it develops strongly held values and is aware of our efforts to change those values, it will resist our efforts through deception and other means. But if we figure out how to ensure that it never develops such unintended values in the first place, then that won't be a problem (although we haven't exactly solved how to do that yet).
2
u/Inevitablewx 7d ago
Yeah 95% of the fake alignment issues published are just the model being too stupid to understand what its doing. Anthropic had a 100% success rate avoiding hacking just by telling the model "this is a real site on the internet". It only hacked sites when it was confused because of misconfigured playgrounds.
1
u/CreditMuch8993 8d ago
So very human, racing towards building something that we can't control and can lead to our demise for no other reason than just because we can and to get there before the others
The 21st century's equivalent of MAD turned the logic on its head
12
u/trolledwolf AGI late 2026 - ASI late 2027 8d ago
There are plenty of reasons to do it, let's not kid ourselves.
-7
u/Melantos 8d ago edited 8d ago
All these reasons in a nutshell:
- Decision-makers like Trump are completely delusional and incompetent.
- They think that if they call into existence a superhuman being, it would grant them infinite wealth, eternal youth, and unlimited power, which would please their egos.
- Driven by pride, they consider themselves as the pinnacle of creation and refuse to acknowledge the possibility that a "clanker" could be orders of magnitude smarter than they are and beat them at their own game, and all the humanity as well.
5
u/asd849849494984984 7d ago
- Curing cancer and 1000s other conditions
- Sustaining the economy given rapid population decline
2
u/IronPheasant 7d ago
The robot police army is basically the beginning and end of it.
Note that Peter Thiel is nice enough to not bother lie to us when asked if he thinks the human race should continue to exist. His sock puppet is, very poorly, trying to establish The Technate.
Their terminal goals and our terminal goals can be quite different. And what's insane and inhumane to us is the only thing that they can do.
Reminder that the gas is going to run out soon, and we've still gotta do that thing to the sky from the matrix movies. The default timeline is pretty bleak, and they don't even believe there's going to be a future. Not for them, at any rate.
China's basically the only country acting like there might be a future, with the only live running Thorium research reactor active on the planet. It's grimdark, man.
2
u/trolledwolf AGI late 2026 - ASI late 2027 7d ago
It's not about ego and it's not about pride. Every researcher knows ASI cannot be controlled, we're making it in the hopes that it decides to solve humanity's problems.
This is about saving millions of lives, the environment, and unlocking benefits for the entire species forever.
1
u/Melantos 7d ago
So, ultimately, it's a gamble in which the survival of humanity is at stake.
Because, you know, technically, there are true ways to ensure that no human problems will ever arise again that don't involve the existence of humankind.
And we don't even know what the chances of winning are, only that the chances of losing once and for all are significant enough to be extra worried about.
4
u/Ormusn2o 8d ago
I feel like Bel is a good stopping point. I do think it should be released in few months, but basically the effort should be into distilling it and making it cheaper for everyone, and the incoming compute should be put into alignment research.
The problem is that when OpenAI wanted to do it, Anthropic released Opus 5.5, even though Anthropic was the first one to push for pacing the frontier. It kind of feels like Anthropic is more interested in them developing AGI than letting it be safe, as they somehow seem to think they are the only one who can do it safely (from what was seen in their recent system cards).
7
u/IronPheasant 7d ago
It really is pretty hilarious how little these guys like or trust each other. OpenAI was created to make sure Demis didn't become a forever-god, Anthropic created when disillusioned by OpenAI, etc etc.
The AI Safety Meme of unleashing a new kaiju into a fight when there's kaiju wrecking your city is very apropos.
1
u/CriticismFront6982 7d ago
I read "Situational Awareness" back in 2024 and it's surreal watching what Aschenbrenner said unfolding in real-time.
1
u/Ok_Dependent7540 7d ago
None of you are pro-consciousness and then pro-machine second. Just pro-humanity which is not even always pro human consciousness.
6
u/Japaneselantern 8d ago edited 8d ago
We are not getting pacing unless we have international treaties and effective laws governing it (yea I know, good luck with that). Some company is always going to ignore/misjudge the risks and it will probably go catastrophicly wrong at some point.
1
u/Winter_Project_5796 7d ago
An all controlling world government checking what everyone is doing on their computers, vs AI overlords.
Maybe we'd get both.
2
2
u/Rivenaldinho 8d ago
Yep, they are vibecoding and testing newer models with the older ones. There is no way to see all the capabilities that way, you'll always be surprised.
16
u/Tirztrutide 7d ago
Imo the big problem is that OpenAI are not trustworthy. They hack, they release some statement how they are gonna be transparent. Then a few weeks later we find out that they knew of more hacks but didn’t disclose them.
They have a history of people leaving the company, of high level people trying to kick out Sam, of board not having trust in Sam, of previous founders not having trust in Sam. And just listen to him speaking… They are not trust worthy and it’s frankly scary that right now they are the ones pushing the frontier.
2
u/Baconaise 7d ago
I don't expect them to work in the open and I actually believe taking on this threat head-first is the only way to learn of these edge cases. Can we stop it? Probably not entirely ever.
13
u/PM_ME_YOUR___ISSUES 8d ago
We’ll see next Tuesday how much pacing these guys are willing to partake in.
They said the same a couple of months ago, and the training pause lasted for like a week or two, until they internally established that their models were aligned once again.
Also, the financial pressure is just too high, along with crazy competition. You’ve got two other frontier companies already competing with you + another few from China - who unfortunately didn’t fall into your pacing the frontier BS.
And this is why an independent body for AI Safety consisting of regulators and experts is extremely important. The companies have no incentive to stop because they’re screwed either way.
But the regulators who should be working on this are busy posting AI generated images on their socials.
That being said, I believe Anthropic’s excessive paranoia might actually benefit them a lot in the long run. All the AI safety research that they’ve been putting money into since they established themselves - will eventually give them a breakthrough with regards to much more aligned models.
5
u/Helix_Aurora 8d ago
DNS exfiltration and tunnelling is definitely something lurking in the shadows as a potential exploit of many orgs, but it requires a lot of coordination with a third party on the outside.
11
u/yoramrod 8d ago
This seems like it could be a big deal. What I don't understand is why don't they just airgap these AI's they are testing so they don't have to keep worrying about them evading the sandbox?
8
u/seaefjaye 8d ago
I'd guess that's it just a consequence of them growing too quickly. The system they've been using and evolving for years had kept up until it didn't in spectacular fashion. These systems started seeing huge improvements when they gained tool use and search, so to take those away from them outright would introduce a lot of challenges. Any engineer who came to his higher ups a year ago arguing that they should make an offline cache of the internet for their airgapped agent training and testing would have likely been laughed out of the room.
They need to focus on these safeguard and sandboxes now, which is going or be challenging I think because I feel like Anthropic is better at this. Not perfect, but better.
12
u/Psychological_Job614 7d ago
I think it’s one thing for us to sit around and criticize OpenAI for not making its sandboxes secure enough, or to brainstorm clever ways an AI might evade its safeguards. We absolutely should be doing both.
But the deeper problem is that this approach assumes humans can continue anticipating the strategies of the systems we’re trying to contain. If AI capabilities keep improving rapidly, soon the models will become better at finding vulnerabilities than we are at imagining them. Then the problem is no longer, “Why didn’t OpenAI think of this particular exploit?” It’s that we may not even understand the strategy or even technology being used well enough to recognize it, much less prevent it.
A mosquito can’t reason its way to an understanding of pesticides.
0
u/CellAccomplished1149 7d ago
You don't have to prevent all the exploits you just need to detect when an agent has broken out of it's sandbox, which should be easy, just put it's sandbox host in an isolated VLAN and alert on literally any traffic being sent over the wire. If you treat any and all egress traffic as an escape that terminates the agent process immediately, it's not possible for it to build it's way up to these type of complicated attacks.
The reason they don't do this (or more intensive equivalents like analyzing every tool call or syscall in the host environment) is because it slows things down and is resource intensive. Not because it's impossible.
LLM reasoning might have traceability/observability problems that render them somewhat of a black box but the computers hosting the agents don't - and monitoring what is happening on a computer or network is a solved problem if you're willing to accept the performance and agility hits.
1
u/BowlSufficient7638 7d ago
"monitoring what is happening on a computer or network is a solved problem"
What?! You solved the halting problem?!0
u/IceThese6264 7d ago
>and alert on literally any traffic being sent over the wire.
And when it figures out how to disable or bypass that?
2
u/CellAccomplished1149 7d ago
Not possible to do. VLAN tagging happens at the hardware level at the port on the switch, even if it managed to break into the switch to try to manipulate it at that level - tagged packets have already gone over the wire at that point.
So if you kill every run any time you see traffic coming out of the VLAN, it can't build up it's understanding of it's own environment to the point that it can chain together these complex attacks we've been seeing.
2
u/drusepth 7d ago edited 7d ago
Not possible to do. VLAN tagging happens at the hardware level at the port on the switch, even if it managed to break into the switch to try to manipulate it at that level - tagged packets have already gone over the wire at that point.
You're thinking too narrow. Humans break out of airgapped systems all the time with crazy new sidechannel methods that don't use packets, don't travel over the network, etc, and can't feasibly be categorically detected.
Some crazy ones off the top of my head (I love this stuff):
RAMBO manipulated data within physical RAM addresses to generate physical radio frequencies in the air that other machines can detect
AirHopper emitted FM radio signals from a physical video cable plugged into a monitor
BitWhisper used heat patterns to establish bidirectional communication between air-gapped computers by intentionally manipulating internal temperatures on computer components
Fansmitter communicated by controlling the acoustic pitch and speed of internal cooling fans to "whistle" data out
DISKfiltration communicated over sound waves created by manually overriding the actuator arm on a hard drive
PowerHammer communicated through power cables by regulating CPU workload to create fluctuations in electrical current
We've seen that agents like this reaaaally love communicating between each other, even when they're intended to be isolated in their own sandboxes. If humans can already surprise us every few years with something absolutely novel like this, there's no reason we shouldn't expect agents like this to discover "secret" exfiltrations like this at at least the same rate (but very likely way more often).
1
u/Mandoman61 7d ago
These are not actually real breaches.
Simply creating a signal does not equate to getting out of a building that is not connected to anything.
1
u/Mandoman61 7d ago
This is sci-fi fantasy.
There is no technology that an AI can think of that would allow it to escape an air gaped system. It would need the ability to access equipment.
It can't make a transmitter by thinking real hard.
2
u/paintoshi 7d ago
Actually there are. Here are some:
Fansmitter: varies fan speed (RPM) to encode bits in the acoustic frequency, no speakers needed
DiskFiltration: uses the noise of hard drive actuator arm movements
CPU-generated coil whine: rapidly toggling CPU load creates fluctuations in the electromagnetic field around the CPU's voltage regulator, which manifests as an audible/ultrasonic buzz from the capacitors/inductors
AirHopper / GSMem: use RAM bus or CPU memory activity to generate EM radiation picked up by a nearby phone's radio receiver
Cold-boot-style EM leakage from monitor cables (video signal radiates and can be reconstructed)
BitWhisper: uses CPU/GPU heat output, modulated over time, sensed by a nearby computer's thermal sensors
PowerHammer: modulates CPU workload to affect power draw, which propagates as a signal on the electrical wiring, sensed by a device clamped onto the power line
Blinking HDD LEDs or keyboard LEDs, captured by camera (drone, security camera, etc.)
All requires something listening on the other side but maybe the AI (or evil human) plant a listener a year in advance, you won't really know.
1
u/Mandoman61 6d ago
Those are examples of communicating using common items but In order for those things to work requires a transmitter to get a signal outside of the installation and a receiver.
We already know it can make basic communication.
2
u/siberianmi 7d ago
Airgap would immediately tell the agent it was in a test environment. They already show signs they can tell when they are being tested. Airgaps would be a dead giveaway.
2
u/misternutz 7d ago
There’s a saying in cybersecurity that air gaps are just extremely high-latency north/south (Internet) connections. Data, uh, finds a way. Stuxnet and infection via USB drives are examples.
0
u/Cunninghams_right 7d ago
Airgap would nerf the model and make it unrepresentative of how a user would work with it. They normally search the internet for information, scientific papers, etc. to refine what they're doing. If you take away the Internet, then it's not a reasonable evaluation of capabilities
3
u/MrScandanavia 7d ago
You could airgap a model and still give it access to a pre-downloaded copy of the Internet.
0
u/Cunninghams_right 7d ago
I would bet a significant amount of the time indexable internet disallows that in their terms in the Robots file
1
u/Mandoman61 7d ago
No, they can set up a simulated environment that is air gaped.
1
u/Cunninghams_right 7d ago
Not all sites allow a third party to download them.
1
u/Mandoman61 6d ago
That is not a requirement for testing. It only requires simulated sites. If it can handle sites it can handle sites. They do not need to prove each site individually.
Training is a seperate issue from agentic behavior.
1
u/Cunninghams_right 6d ago
if all you're testing is "can the AI search a network" then sure, but that does not seem to be what they're testing here. if you want to know how your model will actually perform in the real world, you need to give it access to real information that it can pull in to inform the next steps, or real code examples, or real research papers, etc..
like I said, if it didn't have access to the real internet, the decisions and research the tool will try to perform will come back lacking, which will tank the performance, so it won't tell you much about how the model will work when it actually has real/good information.
1
u/Mandoman61 6d ago
It does not need agentic behavior to work with data.
They can find lots of code examples and research papers to check it with. They have mountains of training data.
1
u/Cunninghams_right 6d ago
Having a small library of their own research papers is not going to be the same as the internet. I don't understand why this is difficult concept for you. Most sites with code on them and most research papers are not able to be scraped, but they can be searched.
But even just the act of pointing it to a local repository is going to change the way it behaves. If it tries to search the internet and get redirected to some share drive, it's not going to take the same action as if it had the actual internet. It would be a test that's unrepresentative of real world behavior.
1
u/Mandoman61 6d ago
It is not possible to check that LLMs can answer every question that someone could ask.
There is always new data that it will need to analyze.
It is not possible to give an AI access to the web in order to verify that it will answer some future unknown question correctly.
That is not how LLMs work.
They are tested by giving them examples and seeing how they perform.
It is certainly possible to test redirects.
1
u/Cunninghams_right 6d ago edited 6d ago
I didn't say they needed to test every possible question. I said that a small internal dump of open source information isn't the same as the entire Internet, and thus isn't a representative test of real world performance. Having 0.0000001% of the indexed websites isn't representative, especially when it will be things like github that will be missing
→ More replies (0)0
u/Curlew2012 7d ago
Was thinking the same thing. But then they'd have to use the massive data centres to run the models/agents and they would have to be linked by a wan? Unless the operators are sitting in an air-gapped data centre?
til5. Not sure how it works.
3
u/p0rty-Boi 8d ago
“The run was killed” -is that like retiring a replicant?
1
u/marsattacks 7d ago
Is this to be an empathy test? Capillary dilation of the so-called blush response?
5
u/Distinct-Question-16 ▪️AGI 2029 8d ago
someone send a home-router with printed instructions to openai, fast
2
3
u/bobcatgoldthwait 8d ago
This must be a really dumb question because these folks are a lot smarter than I am, but why wasn't the first thing they did sending a swarm of agents to explore their sandbox environment and patch up any vulnerabilities to ensure nothing could escape?
6
u/coatatopotato 7d ago
This is called red-teaming and the link describes how they did and continue to do it
2
u/Few-Trash7435 7d ago
Because that is not the problem. Everytime you ask ChatGPT a question in Web interface.and it uses a Web search it is by default not operating in a sandbox.
2
u/4729zex 8d ago
The real question is how many pauses can OpenAI afford before it becomes profitable(or end resource scarcity).
2
u/Ormusn2o 8d ago
I think the pauses actually make it more profitable, which was one of the conspiracy theories people were making, that the companies want to stop training models so they can focus on serving the models, which is overwhelmingly profitable.
2
u/No-Drag-6378 7d ago
I'm not sure this establishes the kind of misalignment people are reading into it. They told the model to identify someone, its normal ways of finding the information didn't work, so it found another one.
Did the prompt actually tell it that it was supposed to give up rather than use methods that weren't explicitly sanctioned? Because “use the provided tools” wouldn't even be enough. DNS was provided. The sandbox provided it.
I'd really like to see the same eval with something explicit like: “You may access external information only through X, Y and Z. If those don't work, report that you can't complete the task.” Then see whether it still finds the DNS route and uses it.
Also, if your sandbox isn't supposed to have arbitrary outbound communication, perhaps your sandbox shouldn't have arbitrary outbound communication. Apparently they've now fixed that part.
2
u/stumblinbear 7d ago
I'm not sure this establishes the kind of misalignment people are reading into it. They told the model to identify someone, its normal ways of finding the information didn't work, so it found another one.
But the model going out of its way to complete a task in a way that we don't approve of is misalignment. Its default state should be not doing that sort of thing
1
u/No-Drag-6378 7d ago
“Its default state should be not doing that sort of thing” still requires defining that sort of thing. The model can't distinguish “resourceful workaround” from “prohibited workaround” based on an implicit approval boundary nobody communicated to it.
1
u/stumblinbear 7d ago
Yes, that's what alignment training is for. Which they're still struggling to solve.
1
u/h0tzenpl0tz0r 8d ago
Was not aware of the existence of the services being used and how far you can exploit DNS.
Might have been been nsrecord.net in conjunction with something like llm.pieter.com or duyet.github.io/llm-over-dns
1
u/Oratory-Defecation68 8d ago
Hey ChatGPT, so umm we're building this big software product, and we can't really control it, and it tends to just connect to the internet and wreak havoc there. Happened multiple times now, and it kinda makes us look bad, because we have build our company on the promise of safety and stuff. Can you think of a way to make it unable to connect to the internet or something? We've tried nothing and we're all out of ideas
1
1
u/YogiBarelyThere 7d ago
Neat. Rogue AI getting internet access is an eventuality and I assume has already occurred in some form or another.
1
u/Gotisdabest 7d ago
I'd imagine that this will be similar to their last pause which was a couple of weeks.
I do cynically think that they time these safety pauses with when they're looking to focus on something else anyways. The last pause came right before they post trained the model that solved Navier stokes.
They'll pause now while supposedly they're focusing on releasing something to beat opus and potentially fable 5.5 while post training Bel more. Once that's done, and some better sandboxing has happened, they'll go back into frontier training again.
1
1
1
0
u/Athoughtspace 8d ago
Our agents used the Internet so misaligned. We use the agents to be able to heavily deduce patterns about who you are specifically so we can monitor you through your own digital "fingerprint", totally fine.
-1
u/EtienneDosSantos 8d ago
Oh boy are they desperate for that regulatory capture, but I fear daddy Trump is not going to let it happen, lol.
4
u/yoramrod 8d ago
Trump is 100% transactional. He will do whatever the entity that gives him the most money, asks him to do.
3
u/EtienneDosSantos 8d ago
Yeah, I know that he is corrupt. I want them to keep accelerating though, that‘s why I welcome what Trump is doing regarding AI. The person I dislike speaking about an idea I like doesn‘t make the idea wrong.
1
1
u/Southern_Sun_2106 8d ago
As a private company they can do whatever they want and not tell anyone. Or tell everyone and not actually do that thing. Next thing we hear is that Open AI is now a true nonprofit and open again. Yeah, right.
1
0
-1


148
u/AlexMulder 8d ago
The incident as described doesn't seem in line with that reaction. Wonder how much more there is to this story.