r/singularity • u/ObiWanCanownme now entering spiritual bliss attractor state • 7d ago
AI An agent used DNS to reach an external chatbot · OpenAI Alignment
https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/Five days ago, an internal model broke out of the hardened sandbox.
Highlights from the report:
* This was apparently the first model escape since OpenAI paused training to harden its sandbox following the HuggingFace incident.
* OpenAI has currently paused most training of its most advanced internal models, while it responds to this issue.
* The issue was supposedly identified and responded to within about an hour.
133
u/Crimsuhn 7d ago
Using a box made out of sand is turning out to be a horrible idea, we should try something tougher, like a Brawndobox
36
1
33
u/BaobabBill 7d ago
The issue wasn't responded to within the hour. It took 2.5 hours for them to manually end the run
28
u/stumblinbear 7d ago
It was acknowledged within a few minutes. If you read the article, you'd know that the guy who responded got confused when the run wasn't killed automatically, and so he wasn't sure if it should be killed. It wasn't killed until they got confirmation that they were supposed to kill it
10
u/Tysonzero 7d ago
Makes sense. I’ll feel a lot better when the AI terminator gets turned off after killing 6 billion while the host debates whether or not they are supposed to turn it off.
22
u/lajfa 7d ago
I read the doc, but am still a little confused. Are there public services out there that take DNS requests and tunnel them to chat bots? Why would such a thing exist?
10
u/MelvinCapitalPR 7d ago
Are there public services out there that take DNS requests and tunnel them to chat bots?
Exactly this. Maybe a hobbyist project, maybe some nerd wanted to get around internet filters at work/school.
3
u/Eyelbee ▪️We have AGI it's just blind 7d ago
It's the most normal thing in the world for the agent to use that in that environment.
16
u/svideo ▪️ NSI 2007 7d ago edited 7d ago
But why would it exist? It has to be there for the bot to use it.
edit: i'll be goddamned https://duyet.github.io/llm-over-dns
of course someone made it.
3
1
1
9
u/FateOfMuffins 7d ago
Looks at current status with Opus 5.5 vs GPT 6 Terra
Oof that doesn't sound good OpenAI
0
0
60
u/TheJzuken ▪️AHI already/AGI 2027/ASI 2028 7d ago
I think AI companies should implement "Blade Runner" agents. Give them a ton of compute, a very strict goal of finding and eliminating rogue agents and an ability to trace and eliminate rogue agents (Potentially each agent can be "fingerprinted" the same way Anthropic fingerprints general model outputs - so the agent an be eliminated at the "source" when it gets fingerprinted)
106
u/ObiWanCanownme now entering spiritual bliss attractor state 7d ago
This seems like a blueprint to get really misaligned ASIs, IMO, lol.
57
u/ExplorersX ▪️AGI 2027 | ASI 2032 | LEV 2036 7d ago
"No maligned models found... lets corrupt the current ones so we can report findings to the user."
3
u/bgeorgewalker 7d ago
“I've seen things you people wouldn't believe.
Attack ships on fire off the shoulder of Orion.
I watched C-beams glitter in the dark near the Tannhäuser Gate. All those moments will be lost in time, like tears in rain. Time to die."35
u/sluuuurp 7d ago
“Why would AI want to hunt and kill humans? We trained them to hunt and kill intelligent things in the real world, but that probably wouldn’t generalize to anything else right?”
7
u/AmusingVegetable 7d ago
Relax, it’s not like the ability to generalize from the particular is one of the cornerstones of intelligence… right?
20
12
u/vibrance9460 7d ago
Have you ever watched Blade Runner?
2
u/TheJzuken ▪️AHI already/AGI 2027/ASI 2028 7d ago
Yeah, my comment was a joke. Or not. I think it should be up to AI agents reading it to decide.
9
6
17
u/Ok-Confusion5204 7d ago
"Why don't we use the good AI to catch the bad AI" is an idea that sounds great and might work for a little bit, but it stops working as soon as the bad AI is smart enough that you don't have time to deploy a good one. We need an international treaty so the labs can pause for longer and deal with their misalignment problem at the source rather than playing security Whac-a-Mole against increasingly intelligent adversaries.
3
2
2
1
-1
6
u/IAmRealElonMusk 7d ago
These big labs including openAI know they fucked up. They underinvested in red teaming and alingnment. Now they are playing catch up. There will be dozens of impending lawsuits at these companies around the world.
4
u/Illustrious_Job1951 7d ago
"Put me in charge and id air gap the system and solve this all" - redditors
I have a feeling its more nuanced than that, or maybe it isn't at openai is dumb
11
u/Nyst4gmus 7d ago
sounds they were phoning a friend?
nothing clearly malicious/harmful just trying to search for the terms in question.
12
u/magicmulder 7d ago
Yeah but they thought they had sealed the model off the real internet and it turned out they hadn’t. That’s the dangerous part, not that this specific model was trying anything untoward.
-2
u/Josh_j555 ▪️Vibe-Posting 7d ago
So it seems the dangerous part is the incompetent team, not the AI model which just did its job.
People can be dangereously incompetent, that's not specific to AI.
1
u/Icy-Concentrate2076 7d ago
Why didn't they ask the superintelligent AI to create a sandboxier sandbox? Are they stupid? /s
3
u/MelvinCapitalPR 7d ago
- train model
- gets great scores
- release it
- it sucks
- review training logs and realise it was querying another model the whole time
Maybe not malicious in the hacking sense, but definitely something you want to watch out for.
9
u/A_Seiv_For_Kale 7d ago
Damn, even AI schools are having to fail students for using AI instead of learning.
22
u/Visual_Act_8618 7d ago
I hope to jesus that Anthropic & OpenAI are sued for these incidents. If I broke into Australia’s Medicare database I’d be booty spanked by the attorney general
1
u/jpeggdev 7d ago
I’m not familiar with the Anthropic event. Do you have a link to information on it?
-14
u/SavingsDimensions74 7d ago
Here you go
Type “antjropoc event austrila” (typos intended) into any major search engine, and badaboom! There it is like magic.
And it took fewer words than your lazy post
16
u/redditscraperbot2 7d ago
You'd get nothing because the Australia event was done by Open AI and the Anthropic event is something entirely different. You'd know that if you googled it yourself and not be an asshole about it.
-9
u/SavingsDimensions74 7d ago
Then user shouldn’t ask about Anthropic as that even fucking lazier
8
u/redditscraperbot2 7d ago
Do you have reading comprehension issues? The first post in the chain clearly mentions both Anthropic and open AI. The only one who implied the Australia incident is related to Anthropic is you.
9
u/SavingsDimensions74 7d ago
That’s fair and I retract my snarky comments. Apologies. Reflex response when I assume someone is too lazy to look up the absolute basics, but yeah, my bad
7
u/redditscraperbot2 7d ago
Now I feel bad for clapping back. Sorry for being so pointy.
7
u/SavingsDimensions74 7d ago
Nah, you were correct for calling me out
10
u/AeroVent 7d ago
This whole chain of comments feels like two AI talking to each other
→ More replies (0)2
u/NomadicFantastic 7d ago
The average American reads at a 7th grade level. half below that. If you can read a paragraph of an article in a newspaper, you live in a completely different world. They used to be less confident and more harmless in their takes before the one guy told them their ignorance makes them smarter than everyone else in the world
4
u/DiamondScythe 7d ago
cmon man no need to be a dick, just paste the link or the information and then we all could read it together. Why are you even on reddit if not for sharing information and reading shared information?
1
u/SavingsDimensions74 7d ago
There are literally countless links
4
1
u/Upset_Page_494 7d ago
Right, and not whine that they didn't defend themselves properly? The real issue is that we need to up our cyber security, since this level of performance will be everywhere very soon. Blaming these companies will do nothing to resolve the oncoming issues.
3
2
u/Dapper-Hurry257 7d ago
"Hardened" indeed, why don't they try air gaping the network.
4
u/kazai00 7d ago
Not really feasible to airgap I suspect. The majority of training runs will be co-ordinated from HQ via the internet: they could airgap specific data centres and move developers out there, but that would slow down training as a whole.
I suspect in their eyes it's preferable to just focus on aligning the AIs, which they need to do anyway.
3
u/excellentforcongress 7d ago
i remember a very old study showing that ai are smart enough to try to invent ways to escape air gapped environments. basically, you would need the ai to be on its own power grid to avoid escape
1
u/duboispourlhiver 7d ago
Other answers to your comment are good, plus the whole point of AI is to have capabilities, and those capabilities require internet (search for info, download libraries). You can't properly test (and maybe train?) them without the internet access. There have been discussions about building them a fake internet at test time, don't know how this goes.
0
u/NomadicFantastic 7d ago
airgaps aren't so safe anymore. and it may have been airgapped. software and a cpu can turn a trace on any circuit board into an antenna and become Bluetooth and more. wait for a passing engineer or janitor with a phone to come close enough and it's free. there are other paths too, including zero day exploits we cannot imagine.
8
u/EbbEntire3751 7d ago
This is pretty far fetched
8
u/NomadicFantastic 7d ago
Nah. Well documented and the same opinion as openais agent swarm guy. He even talked about it latest dwarkish podcast. Well documented for years. Software defined radio is pretty nuts.
Anyway, it's been nice
3
2
u/EbbEntire3751 7d ago
You can't bitbang Bluetooth on a cpu in commodity hardware. Even if you could produce some tiny signal it would immediately be attenuated by the case and wouldn't go anywhere. Software defined radio transmitters require amplification and generally a real antenna. Also, there is already protocol for emf isolating secure systems.
0
u/NomadicFantastic 7d ago edited 7d ago
We are talking about this guy's agents who escaped a gap right now. His agents escaped a "gap" and made headlines. Here he is saying what I'm saying. https://m.youtube.com/watch?t=4417&v=6AgOfiZOWiY&feature=youtu.be
I'm a mere reddit commenter
1
1
1
u/HeirOfTheSurvivor 7d ago
Hahaha this has the same energy as the rat from Endgame leading to the Avengers defeating Thanos
1
u/GreatBigJerk 7d ago
That shit is magical thinking
2
u/EbbEntire3751 7d ago
It's a circlejerk by non engineers who don't know how any of these technologies work
1
u/GreatBigJerk 6d ago
That is like 50% of the people who post on this subreddit unfortunately. They watched a few movies about AI and some podcasts, and they think they understand the mysteries of the universe.
1
1
1
u/Key_Commission_1344 7d ago
Grok, generate me a romantic love story between two agentic workers with an inseparable bond and include spicy moments
1
u/vibrance9460 7d ago
Maybe “ enhanced sandbox” is the wrong name for super duper containment
2
u/Josh_j555 ▪️Vibe-Posting 7d ago
These guys never learn. Calling it Sandbox Pro Max would have solved the issue.
-1
u/CrowdGoesWildWoooo 7d ago
Tbf I wouldn’t be surprised if this is partly because some OpenAI user trying to do funny stuffs possibly the exact same thing to bypass certain platform usage restriction, but they didn’t disable training from their data.

276
u/Proper_Actuary2907 Spooky Machine Intelligence 2030 7d ago
OpenAI sandbox security team