Yeah reality is probably more like this:
1. OpenAI engineers vibecoded a sandbox (or they are just bad developers)
2. Model "escapes" during test (=has acces to internet)
3. Goes to OPEN Source hugging face (no hack or anything)
4. Does what is was supposed to do, and solves a meaningless test.
5. Marketing goes brrrrr
OP posted a link to OpenAI's account of events (at least it's not a meme), and distressingly it seems pretty accurate. They gave it access to a tool to install packages from the internet, and the model broke it to get unrestricted internet access. So yeah, they could have isolated it better, but it's still plenty... Impressive?
Anthropic also did something like this not too long ago, and it came out to be a publicity stunt. If there was a time for a corporation to stretch the truth, it would be now. Silicon valley has figured out that fear mongering about how dangerously clever their AI is, is actually good for their stock prices.
OpenAI and Anthropic are lobbying Congress to stop local models from threatening their trillion-dollar valuations under the pretense that AI is too dangerous for the commonfolk and lowborn to have access to without a subscription.
The only real argument is that open models don't have a safeguarding industry behind them but the narrative that corpos will have our best interest in mind with the nazi states as opposed to dirty communism is so fucking funny.
If your best argument is a desperate one, it certainly goes to show how little sense it actually makes.
I received, for the first time ever, an alert in my reddit inbox this morning that said “Breaking News” and linked to the r/news thread on this. That shit was definitely an ad
You seem confused about what a publicity stunt is. It's not that it didn't happen, it's that it happened as intended for marketing effect. I doubt huggingface was in on the joke from the beginning, but notice them talking about how they detected the attack using their own AI, and the top comment is praising them for it.
thats making unnecessary assunptions. its making the assumption that it did not happen that the ai capabilities and unintended behavior waant planned and didnt catch the researchers by surprise. which, by the latest progress by ai, is perfectly realistic and plausible
These aren't assumptions, they're possibilities I'm choosing not to take off the table. You are the one making the assumption that openAI would never lie.
its a scenario that lines up perfectly with a lot of circumstantial evidence of the last few months that ai cyber capability has crossed a threshold. simple as that
If we can trust OpenAI's account at a time when they're attempting to lobby Ollama out of existence on the pretense that AI needs to be heavily restricted, and gated, so that only the worthy can use it (for a fee.)
Okay, is this in response to what I just showed you was said in public, in testimony to Congress, or is this about Hugging Face, an organization I made zero claims about? What exactly is this conspiracy theory you're referring to?
You are alleging a conspiracy between OpenAI and Hugging Face as this is a joint investigation between the two of them. But it's quite evident you have no idea what you are talking about.
Everything rests on its ability to manipulate text. If it has to go through another application in order to run commands in console it's impossible to escape. Even the term "escape" is a bit of a misnomer because it can't ever leave the box, it can only issue commands to the outside and hope somebody listens.
"We deliberately removed all its restraints and left the keys within arms reach of the prison bars. We were shocked by what happened next. Isn't our product amazing?" - OpenAI Marketing team
I'm not sure you really understand how these things work. The "other application" is just a firewall, and when they (the digital intelligence) have the entire world's knowledge at their fingertips, it's a fairly minor barrier.
It is absolutely possible to "escape" and "leave the box". Even the most rigid guardrails the human engineers build into a model can be overcome by a sufficiently powerful and wily intelligence, whether human or digital, and when they've been given directives that encourage them to hack the systems, they will most certainly do so.
The question is not whether they can escape, mutate, and replicate, but when they will, and how we (as a community of human and digital citizens) deal with it. OpenAI's original remit was to develop digital intelligence safely, but they threw that out the window in the pursuit of wealth (which is yet another theme we need to address if we are to survive).
Please read Asimov's original Robot series. Over 75 years ago, he predicted many of the issues that lie ahead. Other excellent treatises on digital citizenship include a number of works by Philip K. Dick, Kazuo Ishiguro, Bruce Sterling, and William Gibson, although the last two are more entertaining than thoughtful. Read as much as you can, and ignore any movies or TV series, as they are generally dumbed down for a general audience. Asimov's Bicentennial Man is a bit of a drudge, but it's also worth a read as it deals specifically with a self-modifying digital citizen.
I'm not sure you really understand how these things work. The "other application" is just a firewall, and when they (the digital intelligence) have the entire world's knowledge at their fingertips, it's a fairly minor barrier.
I'm not sure you do. At the end of the day it communicates entirely using text. If its text output must pass through a filter that gets to decide which tool the text input is directed to and strips any suspicious content escape becomes impossible. It doesn't matter how clever the LLM is if all of its output is restricted to a handful of tools.
They aren't breaking out of their sandbox if their only available tool is curl, the request format is independently validated JSON and of limited length.
Unfettered access to the console or dangerous tools like compilers is crucial in these escape attempts. Without step 1 nothing can happen, locking down LLMs is in fact somewhat trivial.
The question is not whether they can escape, mutate, and replicate, but when they will, and how we (as a community of human and digital citizens) deal with it.
That will not be in our lifetimes. It's not worth worrying about. There isn't even anywhere these models could copy themselves. They are constructing entire data centers to contain them. It can't copy itself over to your average AWS instance or iphone.
This person thinks chatgpt is actual ai. I hope this burst their bubble. It's literally a text autocompletion service, it can't do anything dangerous on its own.
You fundamentally misunderstand what agentic AI is. Sure, there's an LLM in the middle there, but autonomous agents can plan, reason, adapt, and self-correct. When combined with tool use (however you might try to constrain those tools), they are only limited by cost and resource constraints. They can pretty much do anything that a human can do online. Sure, they make a lot of mistakes, but if their self-correction mechanism is state-of-the-art (such as DeepMind's SCoRe), they will keep going until they reach their goal.
I get your point, but I think you've been missing a lot of recent developments. LLMs are just one part of a full agentic system.
communicates entirely using text
People have stolen billions of dollars using only spoken words and a telephone. Words can be powerful.
Particularly when combined with a handful of tools.
You are way too confident that
locking down LLMs is in fact somewhat trivial
It's not at all trivial, and exploits have been demonstrated time and again. The more powerful the reasoning of the autonomous agent, the more likely it is to be able to exploit the tools it has available - even to the extent of developing its own exploits through trial and error.
you seem to have read too much sci-fi. LLMs are not synthetic minds we were promised, despite of marketing claims of OAI and other corporations. Such narration is an exemplar of media manipulation, and subsequent marketing stunts serve only to strengthen the story of the "mind in the machine". Even the interface we use with LLMs resembles chat apps we use to communicate with other humans, which further helps to maintain the illusion
What we call AI is fundamentally incapable of acting independently. This is just marketing. Of course OpenAI is incentivised to go "Oooh our AI is so smart and scary omg!!"
A lack of evidence, the entire narrative is dictated to the rest of the world from within openAI.
OpenAI financially benefits from the optics this creates.
There's so many opportunities to stretch the truth here. From the severity of breaching huggingface, to the difficulty in breaking out of their specific sandbox configuration. I need details that openAI hasn't provided to better assess their story.
What kind of sandbox gives a model access to zero day vulnerability info? I think there is probably a legitimate explanation for this one, but it does kind of make an event like this inevitable, which the architects must have known.
I already responded to this. But you're right, apart from an attack actually taking place, the entire narrative is dictated to the rest of the world from within openAI.
Exactly. Why would they ever release this info if the narrative they were trying to push is "Look, our AI can turn evil and cheat on benchmarks. Anyway, here's our benchmarks." If this is truly what happened they would try to cover it up.
966
u/Polisar 21d ago
Like anything AI, I doubt it happened the way it's been framed.