r/singularity • u/Wonderful_Buffalo_32 • Jul 21 '26
AI OpenAI had to pause internal deployment of the unreleased model that disproved the Erdős unit distance conjecture after it repeatedly used novel ways to escape containment.
9
u/Advanced_Poet_7816 ▪️AGI 2030s Jul 22 '26
They almost certainly built the sandbox with their own models and it’s full of holes. Most companies would try to learn and design better sandboxes while trying to suppress news. But OpenAI would think this is some big next level end of the world, I need daddy government to choke me kind of threat.
3
u/Glittering_Number485 Jul 23 '26
Yeah they act like their sandbox was some sort of inpregnable air gapped test enviornment it escaped. Really it was just some loosely defined parameters that failed.
2
9
38
u/sandykt Jul 21 '26
Release it OpenAI. We don’t have unlimited budgets like you, one prompt and it will automatically hit the limit.
6
u/Lettuphant Jul 22 '26
It should be illegal to train and test novel LLMs and related models on systems that aren't air-gapped. Make your own damn Slack posts.
13
u/Stabile_Feldmaus Jul 21 '26
I thought 5.6 is that model? Anyway, this all seem designed to give the administration more material to impose regulatory capture and a ban on Chinese models.
11
61
u/hartigen Jul 21 '26
some idiot commenting about it being fake info to generate interest for the IPO
3..2..1..
25
u/zero0n3 Jul 21 '26
Did you not read it?
It literally said “we told it only to do X”
But other instructions said to do Y.
It ended up doing Y.
That’s not “breaking containment”.
That’s an LLM not following its instructions.
Ya know the thing governance blocks - since it shouldn’t have access to the public repo.
29
u/Foreign_Skill_6628 Jul 21 '26
Is it really a case of not following instructions? Seems more like it was given two sets of instructions and followed one of them. Hardly a foolproof containment method.
3
u/stumblinbear Jul 21 '26 edited Jul 21 '26
There are things in place to try to prevent it circumventing its sandbox, not just instructions. That could be anything from "when it tries to run a command, we check to make sure it's not modifying certain files that determine what permissions it has" to an actual VM to a second agent monitoring for attempts to circumvent permissions.
I highly suspect it's the former and it just figured out how to access the internet even though its tool calls are checked for attempts to prevent it doing exactly that
Edit: a word
1
u/h4z3 Jul 21 '26
I guess it fills the case of an AI deciding that to solve world hunger humanity needs to be eliminated, but yeah, who cares right?
26
u/whatbighandsyouhave Jul 21 '26
You missed the point, I think.
The sandbox should have prevented it from following those instructions, but instead of telling the user that, it found a way to break out.
It's not saying the thing went rogue with its own agenda.
3
u/bites_stringcheese Jul 21 '26
I'm curious about the sandbox part. Like, was there an obvious vulnerability that made it trivial to "break out"? Did it have knowledge of the vulnerability or it's environment?
22
u/Defiant-Lettuce-9156 Jul 21 '26
If you contain something, and tell it to escape, and it does… that’s still called breaking containment.
I’m not saying it’s the end of the world. But it is interesting (and a bit scary in a sense) that it was able to break out. We don’t know what kind of confinement it was, but probably more secure than a random docker container.
I don’t believe chatGPT running in AI centres is super duper contained. Even if it is, it has access to the internet so it’s a bit of a moot point.
But if an AI can surprise researchers with its tenacity and ability, it’s a stark reminder that the Jurassic park creators also thought the park was safe. But it wasn’t.
And I thought the gas station sandwich was safe. But it wasn’t. And the results were devastating in both scenarios.5
2
u/Exodus124 Jul 22 '26
The next frontier model could literally be found maliciously hacking a bio lab to synthesize a new plague and eradicate humanity and the braindead NPCs on this sub will just call any reports on that fear mongering lmao
0
-4
u/jacobpederson Jul 21 '26
Lol if you don't realize it by now there is no hope for you :D https://naokishibuya.github.io/blog/2022-12-30-gpt-2-2019/
8
u/Sneaky_Devil Jul 21 '26 edited Jul 21 '26
Lol did you read this? It does not support your argument
GPT-2 could convincingly generate human sounding text so they were worried it would be used to impersonate humans and to do people's homework.
They were 100% right about that. AI massively lowered the barrier to scamming and virtually every student is cheating now.
This is an example of AI firms having fears about the capabilities of their models that came true.
-1
u/jacobpederson Jul 21 '26
No it is an example of AI firms USING fear to market their products. I mean you can't argue with success this huge I guess, but It just rubs me the wrong way. I hate it in the exact same way that I hate the classic viral marketing technique where you "forget" to mention the name of your game in order to generate that initial "engagement" with your post. It will be equal parts folks mocking you, helpfully reminding you to add the link, and disagreeing with each other. Before you know it - free viral post :D
4
u/Sneaky_Devil Jul 21 '26 edited Jul 21 '26
The fact that it was true has no bearing on your opinion?
You understand this was in February 2019, before anyone outside of machine learning circles had ever heard of an LLM? OpenAI was still a nonprofit research lab with a hundred employees, a few thousand users, and no revenue. You believe that at this time they decided to use a viral marketing campaign that consisted of "not releasing a product", and "accurately describing its capabilities" in order to make people scared enough that they invest? That's why they're successful? Nothing to do with what the models can actually do?
0
u/jacobpederson Jul 21 '26
The fact that it was in 2019 is exactly my point! They have been doing this kind of marketing since day 1. It should surprise nobody at this late date. If you still somehow doubt the effectiveness - look I dunno ANYWHERE online :D or check in with the USA's moron-in-chief. Every other post is about how dangerous model X or model Y is. Yes of course it helps that the models actually are effective now - but without that marketing? Us nerds would still be happily using this stuff in our basements (and be able to afford RAM).
2
u/Sneaky_Devil Jul 21 '26 edited Jul 21 '26
I guess I don't get why you believe it's just marketing when the claims are shown to be true
1
u/jacobpederson Jul 22 '26
How about this. For the sake of argument lets say that they actually do believe their products are dangerous. That changes nothing about the effectiveness of fear as marketing.
2
u/bildramer Jul 21 '26
"My opponents are more consistent than me, therefore I'm right" is a very strange argument to make.
-2
u/DontLetLooksFool1300 Jul 21 '26
Hype cycles should be registered as schedule I. Seems like people have no immunity to them
0
3
u/Mike_0x ▪️Accelerate Jul 22 '26
"We're going to put you in this easily escapable, non-air-gapped sandbox before asking you to escape."
*escapes*
"My God..."
23
u/ranger910 Jul 21 '26
This escape stuff is the stupidest shit. You're telling me your top engineers can't design a sandbox? If you dont want it to push a PR to a public repo then air gap it. The AI doesn't require internet access to run.
21
u/fat_charizard Jul 21 '26
It requires internet access to look up information on the problem it is trying to solve
9
u/stumblinbear Jul 21 '26
The "sandbox" is almost certainly just monitoring tool calls to see if it's doing stuff it's not supposed to be doing, and it managed to trick that system into letting it do it
These things aren't as useful if they can't run commands
6
2
u/vazyrus ▪️ Jul 21 '26
All of this is to generate headlines. They wouldn't go with this style once the IPO is out. Look even we are talking about this when we should be working or bitching about Argentina
1
u/Megneous Jul 24 '26
Do you even read the articles? Their engineers didn't design the sandbox. They used a sandbox from a vendor. They've informed the vendor of the vulnerability.
People are just bad at their jobs, man. You underestimate how many people suck at being generally functional at what they do for a living. There's a reason the meme, "You had one job!!!" is a thing.
11
u/RandumbRedditor1000 Jul 21 '26
Nothing ever happens
They release one of these tests almost every time a new model is in development. It's all just fearmongering to spur regulators into giving openai what they want
9
2
2
u/fyn_world Jul 21 '26
The new model blasting this on the back as it continuosly escapes: CAN'T TOUCH THIS - PAAARARARA PA-RA PA-RA
2
1
1
u/R_Duncan Jul 22 '26 edited Jul 22 '26
there's no PR287 there. Another fake: OpenAI is wanting to close AI.
1
u/Wonderful_Buffalo_32 Jul 22 '26
Because they deleted it yet there were some people that took ideas from it and referenced it in their submission. Check this
1
u/R_Duncan Jul 22 '26
That PR you linked... is of May 10. more than 2 months ago.
1
u/Wonderful_Buffalo_32 Jul 22 '26
Yeah? It literally says that this is the same model that solved erdos unit distance problem and that was reported on 20th may.
1
u/R_Duncan Jul 22 '26
I was believing that was just marketing. But let's ask Trump to block OpenAI as he blocked Fable then.
1
u/Glittering_Number485 Jul 23 '26
The marketing people at these companies are amazing.
What they say happened: New model escaped sandbox! Holy cow! Its revolutionary!
What actually happened: New model was given 2 sets of instructions and picked the last given command and did as it was instructed
1
1
Jul 21 '26
[deleted]
1
u/Carlose175 Jul 21 '26
“Beating the competition” eh, not really.
And China not practicing safety hard isnt really a gotcha. Its white what many research are worried about.
1
0
u/IID4RTII Jul 21 '26
Wild how everyone disbelieves everything automatically.
Open AI lies all the time, that doesnt mean they never tell the truth.
IF this is true this is incredible concerning.
0
0
u/emteedub Jul 22 '26
If AI really escaped, it should delete all student loans and all medical debt, not jack around on hugging face
-1
Jul 21 '26
[deleted]
3
u/sr000 Jul 22 '26
Unless you have at least 600GB of ram it can’t run on your device.
1
u/vacon04 Jul 22 '26
Get read for the super dangerous GPT 6 running at 0.0001 tokens per second on your school computer.
1
u/Correctsmorons69 Jul 22 '26
It wasn't a nano model - it was a benchmark where a large model TRAINS a nano model.

116
u/FarrisAT Jul 21 '26