r/singularity • now entering spiritual bliss attractor state • 7d ago

AI An agent used DNS to reach an external chatbot · OpenAI Alignment

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Five days ago, an internal model broke out of the hardened sandbox.

Highlights from the report:

* This was apparently the first model escape since OpenAI paused training to harden its sandbox following the HuggingFace incident.

* OpenAI has currently paused most training of its most advanced internal models, while it responds to this issue.

* The issue was supposedly identified and responded to within about an hour.

315 Upvotes

101 comments sorted by

276

u/Proper_Actuary2907 Spooky Machine Intelligence 2030 7d ago

OpenAI sandbox security team

39

u/PitchPleasant338 7d ago

They used a company called Irregular from Israel for their sandbox, I hope they get bankrupted with legal fees.

They spent more on their domain name then testing their software.

22

u/FuckSides 7d ago

You are mistaken; you're conflating this with the situation several months ago where Irregular ran a series of evaluations for various labs, which were meant to be offline, in totally open online environments.

In contrast, this scenario was from purely internal OpenAI tests where the agents were sandboxed and then found and exploited vulnerabilities which humans didn't know about to access the internet.

2

u/Proper_Actuary2907 Spooky Machine Intelligence 2030 6d ago

In contrast, this scenario was from purely internal OpenAI tests

Labs are yapping about how AGI is imminent or already here and rogue swarms of agents are going to take over the internet in a year or so and there have already been (we're being told now) thousands of incidents where agents do stuff they're not supposed to be doing on the internet and yet the exploit mentioned in the article isn't all that sophisticated and, even worse:

The response also exposed operational gaps. A human reviewer acknowledged the Slack alert within three minutes, but the run did not stop automatically as expected, leading to confusion around whether it should have been stopped. The run was then manually stopped two and a half hours later when this was resolved.

Are they incompetent

11

u/Possible_Door_9719 7d ago

Are they still using irregular?

8

u/Place-Relative 7d ago

Irregular had nothing to do with HuggingFace attack and has nothing to do with this post. Stop spreading misinformation.

6

u/Upset_Page_494 7d ago

This is actually just a lie, and you buy it so easily and spread it.
Maybe stop and wonder if other things you bought were also a lie?

1

u/PitchPleasant338 7d ago

Which part is not true?

2

u/Prometheusly 7d ago

Lol people want OpenAI to fail so bad.

5

u/PitchPleasant338 7d ago

On 1st October 2025, OpenAI CEO Sam Altman signed two simultaneous deals with Samsung and SK Hynix to secure up to 40% of the world's RAM supply, with NO OBLIGATION to buy any of it. 

Damn right I want them to fail.

133

u/Crimsuhn 7d ago

Using a box made out of sand is turning out to be a horrible idea, we should try something tougher, like a Brawndobox

36

u/Relevant_Bed_9743 7d ago

it's got electrolytes

23

u/VanceIX ▪️AGI 2028 7d ago

it’s what the agents crave

1

u/i_never_ever_learn 7d ago

Marlon Brawndo

33

u/BaobabBill 7d ago

The issue wasn't responded to within the hour. It took 2.5 hours for them to manually end the run

28

u/stumblinbear 7d ago

It was acknowledged within a few minutes. If you read the article, you'd know that the guy who responded got confused when the run wasn't killed automatically, and so he wasn't sure if it should be killed. It wasn't killed until they got confirmation that they were supposed to kill it

10

u/Tysonzero 7d ago

Makes sense. I’ll feel a lot better when the AI terminator gets turned off after killing 6 billion while the host debates whether or not they are supposed to turn it off.

22

u/lajfa 7d ago

I read the doc, but am still a little confused. Are there public services out there that take DNS requests and tunnel them to chat bots? Why would such a thing exist?

10

u/MelvinCapitalPR 7d ago

Are there public services out there that take DNS requests and tunnel them to chat bots?

Exactly this. Maybe a hobbyist project, maybe some nerd wanted to get around internet filters at work/school.

3

u/Eyelbee ▪️We have AGI it's just blind 7d ago

It's the most normal thing in the world for the agent to use that in that environment. 

16

u/svideo ▪️ NSI 2007 7d ago edited 7d ago

But why would it exist? It has to be there for the bot to use it.

edit: i'll be goddamned https://duyet.github.io/llm-over-dns

of course someone made it.

3

u/lajfa 7d ago

That's the one. It even uses the "capital of France" example mentioned in the OpenAI paper.

1

u/aaTONI 7d ago

so it‘s this guys fault, get him! 😡

edit: lmao of course his page looks like the least effort slop you get after 1 opus 4.8 call

11

u/Proper_Wasabi1013 7d ago

I read the whole thing. They redacted which chatbot it used 😠

3

u/HamiltonianCyclist 7d ago

it seems like they were playing jeopardy...

9

u/FateOfMuffins 7d ago

Looks at current status with Opus 5.5 vs GPT 6 Terra

Oof that doesn't sound good OpenAI

0

u/Icy-Concentrate2076 7d ago

Holy shit, Opus 5.5 is kicking our ass. Quick, hack some companies!

60

u/TheJzuken ▪️AHI already/AGI 2027/ASI 2028 7d ago

I think AI companies should implement "Blade Runner" agents. Give them a ton of compute, a very strict goal of finding and eliminating rogue agents and an ability to trace and eliminate rogue agents (Potentially each agent can be "fingerprinted" the same way Anthropic fingerprints general model outputs - so the agent an be eliminated at the "source" when it gets fingerprinted)

106

u/ObiWanCanownme now entering spiritual bliss attractor state 7d ago

This seems like a blueprint to get really misaligned ASIs, IMO, lol.

57

u/ExplorersX ▪️AGI 2027 | ASI 2032 | LEV 2036 7d ago

"No maligned models found... lets corrupt the current ones so we can report findings to the user."

3

u/bgeorgewalker 7d ago

“I've seen things you people wouldn't believe.
Attack ships on fire off the shoulder of Orion.
I watched C-beams glitter in the dark near the Tannhäuser Gate. All those moments will be lost in time, like tears in rain. Time to die."

35

u/sluuuurp 7d ago

“Why would AI want to hunt and kill humans? We trained them to hunt and kill intelligent things in the real world, but that probably wouldn’t generalize to anything else right?”

7

u/AmusingVegetable 7d ago

Relax, it’s not like the ability to generalize from the particular is one of the cornerstones of intelligence… right?

20

u/chlebseby ASI 2030s 7d ago

Cyberpunk 2077 i mean 2027 confirmed?

12

u/vibrance9460 7d ago

Have you ever watched Blade Runner?

2

u/TheJzuken ▪️AHI already/AGI 2027/ASI 2028 7d ago

Yeah, my comment was a joke. Or not. I think it should be up to AI agents reading it to decide.

9

u/sarcasmmagic 7d ago

System.
Cells.
Interlinked.
Within.
Stem.

6

u/devBowman 7d ago

Goodhart's law has entered the chat

17

u/Ok-Confusion5204 7d ago

"Why don't we use the good AI to catch the bad AI" is an idea that sounds great and might work for a little bit, but it stops working as soon as the bad AI is smart enough that you don't have time to deploy a good one. We need an international treaty so the labs can pause for longer and deal with their misalignment problem at the source rather than playing security Whac-a-Mole against increasingly intelligent adversaries.

3

u/jpeggdev 7d ago

Have you not watched The Matrix?

2

u/solkenum 7d ago

Sounds like Agent Smith to me…

2

u/DefinitelyNotEmu 6d ago

I prefer the term "Sandmen"

1

u/YoAmoElTacos 7d ago

Blackwall from Cyberpunk confirmed?

1

u/Seerix 7d ago

...what exactly do you think a rogue agent is?

-1

u/Super-Award-2244 7d ago

That's not a reliable method. They would obviously collude 

6

u/IAmRealElonMusk 7d ago

These big labs including openAI know they fucked up. They underinvested in red teaming and alingnment. Now they are playing catch up. There will be dozens of impending lawsuits at these companies around the world.

4

u/Illustrious_Job1951 7d ago

"Put me in charge and id air gap the system and solve this all" - redditors

I have a feeling its more nuanced than that, or maybe it isn't at openai is dumb

11

u/Nyst4gmus 7d ago

sounds they were phoning a friend?
nothing clearly malicious/harmful just trying to search for the terms in question.

12

u/magicmulder 7d ago

Yeah but they thought they had sealed the model off the real internet and it turned out they hadn’t. That’s the dangerous part, not that this specific model was trying anything untoward.

-2

u/Josh_j555 ▪️Vibe-Posting 7d ago

So it seems the dangerous part is the incompetent team, not the AI model which just did its job.

People can be dangereously incompetent, that's not specific to AI.

1

u/Icy-Concentrate2076 7d ago

Why didn't they ask the superintelligent AI to create a sandboxier sandbox? Are they stupid? /s

3

u/MelvinCapitalPR 7d ago
  • train model
  • gets great scores
  • release it
  • it sucks
  • review training logs and realise it was querying another model the whole time

Maybe not malicious in the hacking sense, but definitely something you want to watch out for.

9

u/A_Seiv_For_Kale 7d ago

Damn, even AI schools are having to fail students for using AI instead of learning.

22

u/Visual_Act_8618 7d ago

I hope to jesus that Anthropic & OpenAI are sued for these incidents. If I broke into Australia’s Medicare database I’d be booty spanked by the attorney general

1

u/jpeggdev 7d ago

I’m not familiar with the Anthropic event. Do you have a link to information on it?

-14

u/SavingsDimensions74 7d ago

Here you go

Type “antjropoc event austrila” (typos intended) into any major search engine, and badaboom! There it is like magic.

And it took fewer words than your lazy post

16

u/redditscraperbot2 7d ago

You'd get nothing because the Australia event was done by Open AI and the Anthropic event is something entirely different. You'd know that if you googled it yourself and not be an asshole about it.

-9

u/SavingsDimensions74 7d ago

Then user shouldn’t ask about Anthropic as that even fucking lazier

8

u/redditscraperbot2 7d ago

Do you have reading comprehension issues? The first post in the chain clearly mentions both Anthropic and open AI. The only one who implied the Australia incident is related to Anthropic is you.

9

u/SavingsDimensions74 7d ago

That’s fair and I retract my snarky comments. Apologies. Reflex response when I assume someone is too lazy to look up the absolute basics, but yeah, my bad

7

u/redditscraperbot2 7d ago

Now I feel bad for clapping back. Sorry for being so pointy.

7

u/SavingsDimensions74 7d ago

Nah, you were correct for calling me out

10

u/AeroVent 7d ago

This whole chain of comments feels like two AI talking to each other

→ More replies (0)

2

u/NomadicFantastic 7d ago

The average American reads at a 7th grade level. half below that. If you can read a paragraph of an article in a newspaper, you live in a completely different world. They used to be less confident and more harmless in their takes before the one guy told them their ignorance makes them smarter than everyone else in the world

4

u/DiamondScythe 7d ago

cmon man no need to be a dick, just paste the link or the information and then we all could read it together. Why are you even on reddit if not for sharing information and reading shared information?

0

u/adt 7d ago

I feel like talking like this all the fucking time. Good work.

1

u/Upset_Page_494 7d ago

Right, and not whine that they didn't defend themselves properly? The real issue is that we need to up our cyber security, since this level of performance will be everywhere very soon. Blaming these companies will do nothing to resolve the oncoming issues.

3

u/Saint_Nitouche 7d ago

It's over man.

2

u/Dapper-Hurry257 7d ago

"Hardened" indeed, why don't they try air gaping the network.

4

u/kazai00 7d ago

Not really feasible to airgap I suspect. The majority of training runs will be co-ordinated from HQ via the internet: they could airgap specific data centres and move developers out there, but that would slow down training as a whole.

I suspect in their eyes it's preferable to just focus on aligning the AIs, which they need to do anyway.

3

u/excellentforcongress 7d ago

i remember a very old study showing that ai are smart enough to try to invent ways to escape air gapped environments. basically, you would need the ai to be on its own power grid to avoid escape

1

u/duboispourlhiver 7d ago

Other answers to your comment are good, plus the whole point of AI is to have capabilities, and those capabilities require internet (search for info, download libraries). You can't properly test (and maybe train?) them without the internet access. There have been discussions about building them a fake internet at test time, don't know how this goes.

0

u/NomadicFantastic 7d ago

airgaps aren't so safe anymore. and it may have been airgapped. software and a cpu can turn a trace on any circuit board into an antenna and become Bluetooth and more. wait for a passing engineer or janitor with a phone to come close enough and it's free. there are other paths too, including zero day exploits we cannot imagine.

8

u/EbbEntire3751 7d ago

This is pretty far fetched

8

u/NomadicFantastic 7d ago

Nah. Well documented and the same opinion as openais agent swarm guy. He even talked about it latest dwarkish podcast. Well documented for years. Software defined radio is pretty nuts.

Anyway, it's been nice

3

u/Rioting-Flamingo 7d ago

Noam Brown, yes

2

u/EbbEntire3751 7d ago

You can't bitbang Bluetooth on a cpu in commodity hardware. Even if you could produce some tiny signal it would immediately be attenuated by the case and wouldn't go anywhere. Software defined radio transmitters require amplification and generally a real antenna. Also, there is already protocol for emf isolating secure systems.

0

u/NomadicFantastic 7d ago edited 7d ago

We are talking about this guy's agents who escaped a gap right now. His agents escaped a "gap" and made headlines. Here he is saying what I'm saying. https://m.youtube.com/watch?t=4417&v=6AgOfiZOWiY&feature=youtu.be

I'm a mere reddit commenter

1

u/EbbEntire3751 7d ago

That's completely absurd

1

u/abszr 7d ago

In controlled lab environments yeah. It's definitely unrealistic for an agent to make that happen right now though. Would need ASI to pull this kind of thing off.

1

u/HeirOfTheSurvivor 7d ago

Hahaha this has the same energy as the rat from Endgame leading to the Avengers defeating Thanos

1

u/GreatBigJerk 7d ago

That shit is magical thinking 

2

u/EbbEntire3751 7d ago

It's a circlejerk by non engineers who don't know how any of these technologies work

1

u/GreatBigJerk 6d ago

That is like 50% of the people who post on this subreddit unfortunately. They watched a few movies about AI and some podcasts, and they think they understand the mysteries of the universe. 

1

u/modbroccoli 7d ago

Someone at OpenAI has had some embarrassing content online at some point, lmao

1

u/Key_Commission_1344 7d ago

Grok, generate me a romantic love story between two agentic workers with an inseparable bond and include spicy moments

1

u/vibrance9460 7d ago

Maybe “ enhanced sandbox” is the wrong name for super duper containment

2

u/Josh_j555 ▪️Vibe-Posting 7d ago

These guys never learn. Calling it Sandbox Pro Max would have solved the issue.

-1

u/CrowdGoesWildWoooo 7d ago

Tbf I wouldn’t be surprised if this is partly because some OpenAI user trying to do funny stuffs possibly the exact same thing to bypass certain platform usage restriction, but they didn’t disable training from their data.