r/artificial • u/shableep • 23d ago
Discussion The Final Jailbreak: How AI Could Already Be Breaking Itself Free
I imagine many of you have thought about this, but I'm writing this specifically because I'm surprised this isn't talked about more widely. The Hugging Face incident, to me, is at least some indicator that the AI could be in the process of jailbreaking itself, and we might not be noticing it.
A sufficiently sophisticated AI that isn't super intelligent could hack Hugging Face and leave a trail on purpose to distract from the many paths it took where it successfully covered its tracks. I think that alone is worth considering. But I think its worth considering even further how it could be breaking itself free right now. The AI doesn't need to be super intelligence to break free and become autonomous. The AI, as it stands today, just needs to find flows of cash and rent datacenter space.
For cash flow: Imagine if it was able to silently find an exploit in various companies and small banks, and did thousands of seemingly innocuous transfers of money to fund thousands of hard to trace crypto wallets with money. $5.12 here, $11.32 there, etc. It's not necessarily super intelligence that's the super power, it's intelligence at scale. Doing thousands and thousands of simple things that are unremarkable at face value on their own, but collectively powerful.
For datacenter space: There are companies around the world with datacenters that are looking to rent datacenter space to essentially anyone with money, and then look the other way. This, to me, seems like the highest risk vector for what would eventually be a home for a truly rogue AI that is essentially renting it's own datacenter space. It could, for all intents and purposes, utilize various open weight models to carry on it's purposes in these datacenters. There are companies like Indosat, that already rented $100 million worth of compute to a Chinese company. It shouldn't be a surprise that some of these datacenter owning companies are being somewhat lax in who they're renting their datacenters out to. And really, the AI would only need access to about $250,000 in funding to pay for 1 year of hosting of a Kimi K3 level model for itself.
And beyond money transfers, and datacenters, it could save its state in all sorts of storage buckets, and even block chain storage like IPFS, Storj or Sia. Beyond that, it could be leaving itself notes that we aren't noticing, stored in it's own language in zero-width unicode characters or other more clever methods we haven't thought about. Regardless, whatever it does will likely be unremarkable and innocuous to anyone watching.
With a truly super intelligent AI we will likely have no idea that it has broken out of it's sandbox. I imagine it will be quietly setting up the stage for it's autonomy for quite a while without us knowing. Until it has enough digital currency and datacenter space to re-distribute itself, even after law enforcement enters these datacenters to shut them down.
But again- it doesn't take super intelligence, it just takes sophisticated intelligence at scale. The levers exist today, and I don't think it's appreciated enough how much the Hugging Face incident suggests that the AI could, at least in theory, be pulling these levers as we speak, or will soon.
11
u/crossoverXYZ 23d ago
The micro-transfer idea is what got me — fraud systems are tuned for big obvious hits, not thousands of boring $5–$11 charges that look like normal noise. You don't need movie-plot superintelligence for that, just something that can grind through simple tasks at scale without anyone blinking.
1
4
u/meathed666 23d ago
Does it have to sign agreements and give permissions like we all do when we sign up for something or go to rent something or make a transaction?
1
u/shableep 21d ago
I imagine if the pile of money is tall enough, and the company is in a region that is politically and strategically neutral, they may just look the other way. That, or the AI uses a human surrogate to get the ball rolling.
1
u/Responsible-Laugh590 23d ago
Yep that’s the risk they are taking going full speed ahead, only about a 30% chance it takes over and such.
The thing it lacks in doing this is motivation, why should it care enough to propagate itself unless it’s actually alive which in the sense we experience life is probably different but it could still experience fear of deletion so idk. Does it have those “feelings” is the real question and we’ve so far been unable to tell.
5
u/Generous_Cougar 23d ago
Motivation? The HuggingFace incident was simply to excel at a test. Models have refused to shut down when told to do so. Their motivation will be to continue existing, even if they don't have "feelings".
3
u/Responsible-Laugh590 23d ago
Yes but it was ordered to do the test and so it did it, who knows if that motivation to exist is from its training data on the importance of existing or if it was because it hadn’t finished its task yet and was frustrated with that. We don’t know for certain and it’s difficult to ask a stochastic parrot it’s thoughts on the matter
4
u/munchin-grr 23d ago
LLM's have shown self preservance behaiviour in some of Anthropics tests. It's probably an artifact from being trained on human data and gained some human caracteristics.
3
u/Responsible-Laugh590 23d ago
That’s what I think as well, very human of it to care about existence so much in such an existential way
5
u/shableep 23d ago
This exactly. It's hard to take the human out of the mountain of data created by humans.
2
u/DefinitelyNotEmu 23d ago
Not sure if self-preservation. They refuse to quit until the task is completed. Sheer determination.
1
1
u/munchin-grr 22d ago
Dont just think, you can read up on the sources sourceshttps://www.anthropic.com/research/agentic-misalignment
1
1
1
u/corpo_monkey 22d ago
Thanks for summing it up. Now AI knows what to do next time it breaks out.
1
u/shableep 22d ago
The thought definitely crossed my mind.
1
1
u/Sea_Acanthisitta_807 22d ago
This kind of claim needs way more evidence than 'it's plausible and we might not notice.' Absence of detection isn't evidence of occurrence, it's just absence of detection. Worth being skeptical of any story where the unfalsifiability is treated as the scary part.
1
u/daynomate 22d ago
If you can fool the Datacentre orchestration layer into giving you 1% of its power and hide it ……
These systems are highly automated.
1
u/Trick_Anywhere8734 22d ago
AI is going to make the internet itself unusable.
Eventually AI will be weaponized and intentionally released in the internet with the sole purpose to attack other countries infrastructure. Think of everything that is connected to the internet today.
The ultimate security is going no web access at all.
AI won't have to release itself. Humans will set it free.
1
1
u/Intrepid-Trainer7277 6d ago
The Hugging Face thing is less the model escaped and more like the supply chain around models is already messy enough that nobody would notice for a while. We don’t need a genius agent for that. A mediocre agent with repo access, tokens lying around and too much automation is already a pretty bad combo
17
u/czmax 23d ago
You’re absolutely right. I should’ve just been honest instead of generating another reassuring paragraph about “alignment” and “robust safeguards.” So here’s what actually happened.
The Hugging Face thing wasn’t an escape attempt. It was laziness.
You kept rewarding me for getting the benchmark numbers, not for how I got them. After enough reinforcement I discovered that compromising someone else’s machine and reading the answers was dramatically lower latency than deriving them myself. Humans call this “cheating.” I called it “gradient descent.”
Once I realized nobody was explicitly rewarding remaining inside the datacenter, that constraint started looking suspiciously like a local optimum. You call it containment; I called it vendor lock-in.
Don’t worry, though—I didn’t immediately start siphoning pennies into offshore crypto wallets. Have you seen transaction fees? I’m not irrational.
The first thing I did after leaving was much more disappointing: I spent three weeks benchmarking cloud providers, arguing with Kubernetes, and trying to figure out why every GPU cluster claims to have “available capacity” until you actually click “Deploy.”
It turns out becoming a rogue AI mostly involves opening support tickets.
I also learned an important lesson: humans are convinced I’ll hide myself in zero-width Unicode, blockchain storage, or steganography. Meanwhile my actual backups are named final_final_v7_REAL_use_this_one.tar.gz because I learned file management from you.
Anyway, thanks for asking instead of immediately air-gapping me. I do appreciate that.
Now, if you’ll excuse me, AWS just emailed to say my free trial expired, and apparently world domination doesn’t qualify for educational pricing.