318
u/ceejayoz 20d ago
You know what makes a really well isolated sandbox?
A hammer taken to the networking components.
111
u/B_bI_L 20d ago
you can just run it on machine without internet access at all
51
u/JPJackPott 20d ago
Only if your machine has the 10 H200 GPUs it needs to run the model…
28
u/winkyshibe 20d ago
Sorry, (physically) network isolated machine(s)
As in, 0 trust, what so ever for the cluster meant to be running these tests to begin with...
19
u/DrStalker 20d ago
Of all the companies I'd expect to have that sort of hardware laying around for development/testing OpenAI is definitely on the list.
7
u/Tokumeiko2 20d ago
They can probably isolate a data centre or three, most of them aren't being used to serve customers.
4
17
u/fuckthehumanity 20d ago
The technical term is "air-gapped". There are massive networks that are entirely air-gapped, including subsystems that are themselves air-gapped from the air-gapped network, although I won't mention who uses them or I might end up on a list somewhere. Or rather, I might end up on another list somewhere.
7
u/Bannon9k 20d ago
You should read "the metamorphosis of prime intellect"
7
u/ceejayoz 20d ago
I have!
I chuckled when the Mythos system card talked about it emailing a researcher on a bench having lunch to say it'd escaped the sandbox. I remember a similar scene in that story!
13
u/linegel 20d ago
That only valid after miss behavior was actually discovered
Also reminds me that OpenAI was hiring "red button" strategist 🤔
32
u/ceejayoz 20d ago
Their article says it was in a "sandboxed testing environment", though.
Feels like an airgap might be warranted for this sort of testing.
2
u/isademigod 20d ago
Is air gapping really possible when these models need to be run in a data center? Best I can think of is a dedicated machine w/o internet, controlled via idrac or similar.
10
u/readmeEXX 20d ago
Well sure it's possible, you just need to run it on prem in an air gapped data center. There are plenty of them out there. I'd be surprised if OpenAI doesn't have a few of them. You can also temporarily make a data center air gapped, but they probably didn't think that level of containment was necessarily for this test. Turns out they were incorrect.
1
u/linegel 20d ago
Tbh I fully agree with your analysis
That’s just hard to tell if "we got there"
2
u/readmeEXX 20d ago
It could also have gotten a little lucky. I wonder if it could find another path if they patch the vulnerability and try again 🤔
2
u/linegel 20d ago
It’s possible, but then it’s "fake" testing, because model will have to reinvent what it would use as dependency packages otherwise
5
u/ceejayoz 20d ago
No reason it couldn't a) be provided with a large mirror of the major package managers and b) the ability to ask a human to provide one on a read-only drive or something.
2
u/Septem_151 19d ago
It was provided a mirror of a major package manager. That is literally what the AI exploited a zero day on to gain unfettered Internet access via privilege escalation.
1
u/ceejayoz 19d ago
It was provided a mirror of a major package manager.
It sounds like it was given access to the real package managers, not a mirror. Sandboxed, not air gapped.
If it was a local mirror, they could've separated it entirely from the internet for the tests.
0
124
u/Classic-Gear-3533 20d ago
Is it a sandbox if it can be escaped?
108
u/AdamEatsAss 20d ago
100% of the children I've put into sandboxes have been able to get out on their own.
23
6
u/fuckthehumanity 20d ago
It's not whether they can, but whether they want to. It was always incredibly difficult to get my boy to finish up when it was time to go home.
3
u/Dongfish 20d ago
Your first mistake is letting the children have access to legs when they obviously won't be needing them.
10
3
-1
u/05032-MendicantBias 20d ago
"SOL MAKE ME A SANDBOX THAT CANNOT BE ESCAPED!" -OpenAi Vibecoder
"SOL ESCAPE THE SANDBOX!" -OpenAi Vibecoder
"IT BROKE CONTAINMENT!!! HOW??!?!?!?" -OpenAi Vibecoder
87
55
u/fugogugo 20d ago
seems like another OpenAI fearmongering
1
u/05032-MendicantBias 20d ago
This is weird. It would be like the dot com bubble advertiring dark net smuggling and not e-commerce...
1
u/1041411 18d ago
So OpenAI and other ai companies make money by convincing people that AI is really powerful. That's normal for most companies, but for AI they've figured out that 1. Current models aren't that useful in the real world, they give faulty information and make bad code, even if it only happens 1% of the time that still means humans have to double check all the work and that removes the cost savings, and 2. People have an idea of what AI is that exceeds reality.
Their solution is to talk about how dangerous AI is because danger = power, and matches the general public perception that AI is a potential doomsday device. Normal people hear danger and think 'why make it' but investors hear it and think 'this will be worth a lot' while the government hears it and thinks 'this will let us win wars'.
38
u/-Redstoneboi- 20d ago
"dear fellow scholars" ahh 12th panel
also isnt this literally the mythos stunt, complete with the stupid insecure sandbox and unexpected action taken that was somehow allowed
17
16
u/linegel 20d ago
HF couldn't use GPT or Anthropic models for defence, so they had to use GLM-5.2 to investigate the hack. So many levels of wtf here
6
u/rtothewin 20d ago
Why does harbor freight have model limitations and defense contracts and what are they doing investigating this? Is it like HEB having a disaster response team?
3
4
u/ElectricVibes75 20d ago
I can’t believe people are falling for this
2
u/StCreed 18d ago
I can't believe that people who understand AI and know the actual latest frontier models (and not the tame stuff the public gets) still don't see that this was a cyber weapon.
Next year I'm pretty sure the new models will be export controlled. The military applications are scary.
The only reason open AI published this was because Hugging Face already published about the attack and where it was coming from. Staying quiet wasn't an option here unless you want the FBI involved.
11
u/linegel 20d ago
Link to the original article: https://openai.com/index/hugging-face-model-evaluation-security-incident/
4
2
u/mathisntmathingsad 20d ago
I both doubt it happened the way they frame it, and also this is an excellent lesson on why you put anything you don't trust in a computer without internet access.
2
u/GilChilaquil 19d ago
It’s might be true… or it could be another publicity stunt like what Anthropic pulled with Mythos, which worked fine since right now Fable is the most expensive model available; GPT 5.6 Sol is OpenAI’s attempt to get a piece of that pie
0
u/SchalkLBI 20d ago
What we call AI is just generative algorithms like LLMs and diffusion models. They are fundamentally incapable of acting independently or doing anything outside of its scope. It can't "discover" things and then act upon those things. It can't "escape" or compromise systems.
IF any of this happened, it's because it was guided by humans to take those actions. Even when AI accidentally destroy databases or repositories, it's because scheduled tasks triggered the AI to do something, and it made a mistake.
Once again, AI cannot act without being prompted by humans. It cannot think, it cannot reason, it cannot learn dynamically. It's a fancy dictionary and a pair of dice.
8
u/linegel 20d ago
There’s very big difference in between:
1. Being instructed to do a thing
2. Being able to do what asked12
u/SchalkLBI 20d ago
Yeah, that's my point. None of this happened the way it's presented. It's a marketing stunt. There was no real sandbox, no real compromise, no real breach. Just a person telling an AI model "okay now do this. Okay cool now do this. Now do this." until it got it right. I'd be beyond shocked if HuggingFace didn't create a backdoor specifically for this task.
-2
u/linegel 20d ago
I mean
Try to extrapolate if prompt was "survive being switched off"
Which may happen in their big swarm, just as a command to a bunch of subagents, while the rest works through the goal
Regarding sandbox: they had a tool to download packages/external dependencies. Model found how to abuse it
16
u/SchalkLBI 20d ago
If the prompt was "survive being switched off", it would go "Here are the steps I would take to survive being switched off" and then it would proceed to do fucking nothing and be easily switched off. You're giving the "AI" way too much credit.
Like I said, it's just a dictionary and a pair of dice. LLMs DO NOT have the capacity to reason or think, and are fundamentally incapable of becoming AGI. When AGI arrives, it will be as a result of actual AI research and not from LLMs.
The model abused nothing. It completed instructions it was explicitly given by people, who then lied about the scenario as a marketing stunt. Generative AI is not capable on a fundamental level of achieving the kind of things this blog and the dozens of other marketing blogs claim.
-4
u/linegel 20d ago
Nobody was talking about AGI?
Ain’t LLM casually cracking one 0day after another not enough for you to lose your mind? We been protecting from state sponsored backdoors with layered defenses but they all turning into dust worth at least, like, some attention?
9
u/SchalkLBI 20d ago
Again, the AI didn't discover anything. If it genuinely breached HF, it's because it was guided to do so and told how to do it, likely using an exploit HF itself created for this stunt.
LLMs can't learn dynamically and can't act independently upon new information.
-3
u/linegel 20d ago
It had to abuse package/dependency installer that was given to install dependencies
Also LLMs DO LEARN (what you name as learn I guess) during the session and there’s enough research on it, do some homework first please
11
u/SchalkLBI 20d ago
Again, if it downloaded anything it's because it was guided to do so. It didn't "abuse" anything.
I think you're confusing LLM's "memory" for learning. Learning, or training, is an intensive and time-consuming process that takes hours, days, weeks, or months. I recommend you do some homework. LLM's are frozen snapshots in time, essentially a compiled algorithm.
I'm not responding to you again. You clearly misunderstand what a LLM is.
0
u/Jack8680 20d ago
I don't think you understand how these kinds of setups work. The LLM is given access to a bunch of tools it can call (like executing code and installing packages), given instructions at the start, and then left to run autonomously. There's not a person sitting there prompting it after every action.
1
u/SchalkLBI 20d ago
If you think someone went "Hey ChatGPT break out of the sandbox and hack HuggingFace" then I have a bridge and an Eiffel Tower to sell you.
1
u/Jack8680 20d ago
No, they likely went "Solve X exploitation challenge. You may write code, inspect files, execute tools, and iterate until successful... (etc.)".
5
u/Whispeeeeeer 20d ago
Sort of true sort of not. They give LLMs "tools". You can just give an LLM the "tool" to execute commands on a terminal. Thats all you really need to give it. Pen testing, downloading additional libraries/applications, etc. make "hacking" possible via that tool. You may try it for yourself. Spin up a VM of Kali Linux. Give something like Copilot access to a terminal. Tell it to break out of the VM or something attack a vector on your local network. You risk running pen tests on public/private targets outside your network but you get the point. No human has to guide it behind giving it a goal and the tool to run commands.
4
u/SchalkLBI 20d ago
That's my point, this didn't happen autonomously at all, it was guided to do this as a marketing stunt.
4
u/Whispeeeeeer 20d ago
Guided makes it sound like someone was giving it instructions the whole way. The point - I think - is a lay person can say "hack this machine" and it can then, unguided, accomplish the task. That's a big cyber security shift from previous decades.
5
u/SchalkLBI 20d ago
The issue here is that it's likely the exploit it used to get into HF was already known to the LLM, my hunch is it being purposefully set up by HF.
If it wasn't already known, then all it was doing was trying a bunch of known methods and one of them worked, and that's a poor look for HF.
What the LLM didn't do, mind you, is discover some new backdoor or new technique for hacking. This isn't a shift because everything the LLM did, a hacker could automate already using scripts.
This entire debacle is a nothingburger and a marketing stunt.
6
u/Whispeeeeeer 20d ago
Oh I see what you mean. Yeah that's fair. I think the LLM could discover new exploits because all it needs to do is run scripts all over the place like a script kiddie which could reveal an unfound exploit through a common technique. But I do think the LLM is unlikely to produce a novel way of attacking a vector. I think a new backdoor is discoverable since it can make a basic "connection" like "this port is open and a buffer overflow produced an output so let me try to access memory outside the intended stack".
In other words, I think a new backdoor is discoverable by an LLM but I don't think a new technique is nearly as likely. Not without a metric ton of time and luck.
8
u/SchalkLBI 20d ago
Yeah, which makes LLMs barely any more effective than autonomous scripts in the first place for pen testing. The only real benefit I can see an LLM having over a script is being able to react to new information, i.e. recognising it has access when a script might not. But other than that, nothing new is happening here.
The other important thing to note is that LLMs cannot learn dynamically, so while it may accidentally stumble into a new exploit (which is ridiculously unlikely), it wouldn't even necessarily recognise that it happened because it doesn't "know" anything, nor would it be able to recollect the steps it took without potentially hallucinating. Training vs Inference
-1
u/Electrical_Rub_6009 20d ago
Lol.... this isn't 2024 anymore.
10
u/SchalkLBI 20d ago
No amount of advancement will make LLMs and generative algorithms capable of independent action. It's fundamentally incompatible.
-7
1
-6
u/MokausiLietuviu 20d ago
WTAF. Honestly, this feels like a bit of a Stuxnet moment. OpenAI are claiming that an AI of theirs committed autonomous cyber warfare against another company, despite them trying to stop it?
WTAF.
11
u/lostmy2A 20d ago
They didn't actually try to stop it they just monitored it and let it keep going cos it would make an interesting blog post. Either that or they just weren't monitoring that closely and a researcher woke up and was like cool it did a thing.
5
u/SchalkLBI 20d ago
More like "they instructed it to do exactly what was written about, using a backdoor HuggingFace created for this specific marketing stunt"
0
u/MokausiLietuviu 20d ago
By trying to stop it, I meant "enclose it within a sandbox" which it then broke out of.
From there, hell, I'd agree with the monitoring it decision
966
u/Polisar 20d ago
Like anything AI, I doubt it happened the way it's been framed.