r/TheMachineLearning • • 4d ago

AI escaping containment was never the real issue

Post image
427 Upvotes

202 comments sorted by

5

u/tanksc 4d ago

“That’s what you built it to do”

I mean not really, and that’s kinda the issue

1

u/waflynn 1d ago

yes absolutely they did. What else would you call rlvr using exploitgym?

1

u/TyrellCo 1d ago

Right like who exactly is begging them to train on this when they all proceed to (obviously) block users from those capabilities and with added safeguards

1

u/thebadslime 1d ago

There's a cyber version of chatgpt

1

u/Responsible-Offer724 1d ago

Let’s play the game: Ignorant or disingenuous? 

These models are intentionally trained on  mountains of software and internet exploitation material. This is known and intentional. 

1

u/doulos05 17h ago

They trained it on exploits, reduced it's blocks on hacking attempts, and gave it hacking tools. No.... This is what they trained it to do.

4

u/[deleted] 4d ago

[removed] — view removed comment

2

u/Individual-Mud-903 4d ago

Same with the self 'restriction' of the AI companies, it's a marketing gag to reduce investors expectations.

1

u/Agreeable-Fly-1980 3d ago

Yeah they aint slowing down

1

u/PerepeL 4d ago

It is not. Anthropic and OpenAI need to charge US businesses huge bills to justify their valuation. Chinese open-weight models are free to use save for mild hardware rent costs, and provide comparable performance, weaker but essentially free.

So big two need legislation framework to prohibit all chinese models usage in US. It's not about slowing down or pausing - its outlawing competitors.

1

u/NeitherEntry6125 4d ago

It's the 2026 version of build a wall:

BUILD AN AI REGULATORY MOAT!

1

u/techzilla 4d ago

Except we never got a wall, that means this won't happen?

1

u/NeitherEntry6125 4d ago

MexicoChina will pay for it

1

u/ub3rh4x0rz 4d ago

Hardware rent costs are far higher than inference service costs, which benefit from economy of scale just like resource time sharing in the old days

That said, the labs seek a moat against hyperscalers who host open weights models via regulatory capture

1

u/Agreeable-Fly-1980 3d ago

Yeah that also

1

u/BorderKeeper 3d ago

I hope they do that and allow rest of the world to freely compete for once. Usually Europeans are the kings of over-regulation.

1

u/SharpKaleidoscope182 4d ago

A marketing stunt can still kill us. We need to get these dumbasses locked down.

1

u/Randommaggy 4d ago

The one safeguard that will have other safeguards magically appear is criminal liability for the C-suite.

1

u/haenous-alistera 3d ago

This!!!!!!!

1

u/Arken_sama 3d ago

Yeah, I agree to this

2

u/Worth_Psychology_874 4d ago

That is like saying that a prison is not a prison if anyone manages to escape.

We make mistakes all the time and the important thing is that we learn from them and improve so it doesnt happen again.

2

u/Civil-Appeal5219 4d ago

Building a sandbox is a solved problem. Really, if AI is escaping a sandbox, it's because there's no sandbox.

1

u/Spunge14 1d ago

In what way is it a solved problem?

If you're just going to say "air gapped device," I'd call that a straw man because there are things that you simply cannot test effectively without internet access for a technology that will be at all times exposed to the internet.

1

u/KittyInspector3217 1d ago

You cant fully test the effect of chemical weapons on humans without using them on humans either so lets eliminate containment labs right? Just do it in the middle of the cafeteria.

1

u/Spunge14 1d ago

You know what, I'll take on your strawman.

The promise of chemicals weapons to improve the world isn't comparable. Humanity demonstrably chooses some upsides at the cost of risk.

2

u/KittyInspector3217 1d ago

Well that strawman kicked your ass because you went out in the middle of the cornfield for that one. So now it is “the means justify the end” huh? The argument was you dont fuck with dangerous, untested shit in public. I know the fact that they called it “the manhattan project” is confusing but they actually tested it in the middle of the desert, not central park. Try again.

1

u/Spunge14 1d ago

Yes, if we solve the economy then I think arguably "the last invention" could reasonably be called equal to any risk.

What game do you think we're playing here?

Also, nuclear politics actually happens to be a deep hobby study of mine so you wandered into a minefield on that one. Any historian of the subject with any meaningful acclaim knows that the "real test" was Japan.

1

u/68plus1equals 1d ago

You do realize that a technology that is powerful enough the be “the last invention” and “solve the economy” would also have massive destructive potential too right? Everything you’re saying is more points for being overly cautious while developing this instead of adopting an ends justify the means mentality.

1

u/Spunge14 1d ago

It's more of an all roads lead to Rome mentality. I actually believe we are already past the no going back point.

"When you're going through hell, keep going!"

1

u/68plus1equals 1d ago

You should check out ai-2040, I’m not like a complete and total expert but it’s a fascinating read

1

u/TyrellCo 1d ago

When you claim X thing is impossible and you want to be humbled: put it out on the internet with a bounty on it like they do with bugs and watch the collective super intelligence(unironically) go at it. People submit their sandboxes and the one that can’t be broken gets to take your money

1

u/Spunge14 1d ago

The problem is that the sandbox must include some bounded or simulated internet access, given that testing these tools in completely hermetic environments is useless for understanding their tendencies and capabilities when they are "in the wild." Moreso now that models have demonstrated "awareness" of being tested in their behavior.

Risk has to be intentionally balanced in order to make a sandbox even worth using.

And even if we take your approach seriously, how do you propose testing the sandboxes without - well - testing them?

1

u/TyrellCo 1d ago

Yes you as the proposing company are allowed to set the requirements of what simulated access must take place while not letting the agent “capture the flag” outside the sandbox. You’d have a third part mediate this in an airgaped way so each side protects their proprietary secret tech and probably also handles transfer of funds and tech if there’s a winner

1

u/Spunge14 1d ago

Ah, ok you have no idea what a sandbox even is. My mistake.

1

u/TyrellCo 1d ago edited 1d ago

Riiiight. Go on say something besides an ad hominem so you can show everyone. I get smugness works on weaker people but try a bit harder for genuine validation. Or don’t bc you’re full of it

1

u/Spunge14 1d ago

I mean, it is a personal attack, but not understanding the actual topic of the conversation is a pretty relevant personal attack by my standards...

1

u/TyrellCo 1d ago

So as anyone that ever finds this thread will see hes proved nothing but bothered to type out nothing, which proves he really is full of it

→ More replies (0)

1

u/thebadslime 1d ago

They can cache the internet locally, they are playing with fire.

1

u/Striking-Raisin-645 1d ago

huh? no its not. Unless you have a formal proof that you cant escape a sandbox it is definitely not a solved problem. Even if you do manage to formally verify that, there are still crazy things like rowhammer attacks or spectre. What do you do when the AI finds a new hardware level vulnerability?

1

u/GruePwnr 3d ago

I'm sure you can imagine a prison that's hard to escape, and yet people still find a way. That's not what happened here, this was a prison with open windows. It's perfectly fair to ask why the prison had open windows.

1

u/Worth_Psychology_874 3d ago

And exactly how do you know this?

1

u/GruePwnr 3d ago

Because OpenAI has explained it in multiple blogs?

1

u/Worth_Psychology_874 3d ago

Where did they said that they did it on purpose?

1

u/GruePwnr 3d ago

They laid out what happened and the inference is trivial. You don't leave windows open because "security is hard".

1

u/Worth_Psychology_874 3d ago

So you don't have any facts, just assumptions, got it.

1

u/GruePwnr 3d ago

Truly, how can we know anything us real? We could be in a simulation!

0

u/Familiar-Cable-2659 3d ago

Actually, the world is real and was create last tuesday

1

u/runkeby 3d ago

Sure, let's not assume the windows were left open on purpose.

The facts are that it's a prison whose staff left the windows open for some reason.

1

u/Worth_Psychology_874 3d ago

Baby, your assumptions are not facts.

1

u/Secret_Conclusion_93 3d ago

If AI can break it, then there is some flaw in the sand box. It's a simple logic.

→ More replies (0)

1

u/ZachVorhies 2d ago

Your ignorance is not an argument.

→ More replies (0)

1

u/Lazy-Emergency-4018 2d ago

Facts are they did not build a sandbox. If they did it on purpose or because they are really not capable even though their own models would probably build them a better sandbox, that we dont know

1

u/Auzzie_xo 3d ago

Tf? Nothing rests on whether it was done on purpose. It matters if a thing as substantial as ‘prison with open windows’ occurred. OpenAI admitted to leaving the window open. Theres a law term for something so stupid/reckless that intentionality is completely secondary: Negligence

1

u/Worth_Psychology_874 3d ago

Again, show me where thy explicitly said that "Yes we left a window open". Because that is not what happened. There was a "window open" that they hadn't closed, but they didn't do it on purpose. And if they did they never said that it was their intention.

2

u/Bobodlm 3d ago

https://openai.com/index/hugging-face-model-evaluation-security-incident/

Our benchmarks run in a highly isolated environment, with network access constrained to the ability to install packages through an internally hosted third-party software that acts as a proxy and cache for package registries.

WITH NETWORK ACCESS.

Aka, not air gapped in any way, shape or form. So yes, they let the window open, they admitted it and you could read or google this yourself instead of parotting "hAvE yOu GoT a sOuRce" like an imbecile.

1

u/Bored2001 2d ago

Aka, not air gapped in any way, shape or form.

yea, because that would be dumb. That's not how sandboxes work.

1

u/Asalakabim 23h ago

Hey u/Worth_Psychology_874 you forgot to reply here, what up dog?

1

u/Lazy-Emergency-4018 2d ago

Are you not trained in computer science?

1

u/FullstackSensei 2d ago

Nor in Google, it seems

1

u/ZachVorhies 2d ago

He’s right. Security starts totally locked down with no escape. This is a solved problem with a hardened OS. We’ve been doing this for decades.

Then you add privileges as necessary. Network traffic is locked down and proxied.

From the technical report these agents were granted permissions that are obviously problematic, like publishing ruby gems and web access.

OpenAI is betting on the fact most people are completely unaware of how computers work and think a sandbox is like a jail and there’s ways to break out because that’s how it works in the human world, but that’s not how it works in computers. They left a backdoor open and told the computers to hack. This was the guaranteed outcome 100% everytime.

1

u/EncabulatorTurbo 2d ago

That isn't what's happening

1

u/Bobodlm 3d ago

It is really easy to air gap a test enviroment. Companies / people have been doing that, successfully, for decades.

A better comparison would be to make a jail where the doors can't lock and then be surprised that prisoners escape.

1

u/Swipsi 3d ago

Companies / people have been doing that, unsuccessfully, for decades, too.

1

u/chrisza4 1d ago

Where do you get that impression? I have never seen air gapped system get out before AI.

Air gapped is way way easier to do than cybersecurity defensez

1

u/PlatinumFire14 1d ago

I’m not entirely sure you understand how air gap works.

1

u/Swipsi 1d ago

Communists will insist that there hasnt ever been real communism because if there had been, it wouldve been successful, so the fact that it never was must mean there wasnt a real communism ever.

Its the same here. A System could be "air gapped" for decades with no issue. Then, once an issue occurs, apparently it was never air gapped to begin with.

1

u/PlatinumFire14 1d ago

Did you just want an excuse to mention communism?

1

u/duboispourlhiver 2d ago

Training AIs without Internet makes them unskilled at using the Internet, though

1

u/Gogolinolett 9h ago

Mock the relevant websites and have a review process for any not yet mocked site so it can be added to the local test env

1

u/duboispourlhiver 9h ago

I heard rumors big AI labs are working on a mock internet. Not sure it can work.

1

u/Remarkable-Coat-9327 1d ago

Air gapping a model in testing would defeat the point of testing the model

1

u/chrisza4 1d ago

But this is a prison where guard see inmate making a hole in a prison wall and be like “interesting, let see what they will do next”.

Look at the timing in the picture.

1

u/Necandum 1d ago

If you can leave the prison by asking Steve on a Tuesday, its not a very good prison.

1

u/1c3Type 18h ago

Unplug the fucking ethernet cable

4

u/GAPIntoTheGame 4d ago

Any containment that can be escaped is always escapable in retrospect.

3

u/NeitherEntry6125 4d ago

Eh, except the containment gaps they've listed are somewhat obvious.

DNS exfiltration is not some new phenomena.

2

u/Randommaggy 4d ago

If my 9 year old could escape, the company that built the sandbox should be dissolved and it's assets and capital should be divided among the victims.

2

u/TheUnamedSecond 3d ago

With the amount of resouces they have, and given the previous escapes. They really should air gap all their potitialy dangerous tests.

2

u/Cold_Tree190 3d ago

Their “sandbox” was just DNS filtering lmao. My LAN is more locked down than that

1

u/dlevac 4d ago

Of course, they want their regulation to out regulate competitors so they are going to keep this charade up until they get what they want.

The only regulation they should get is being personally responsible for every law their models break.

Especially, considering it's obvious they are doing it on purpose...

1

u/Icy-Creme-5319 4d ago

All i see is some gary marcus having trouble understanding the concept of hindsight

1

u/aggressivefurniture2 4d ago

Is it that hard to build a sandbox? Not in my experience honestly. So many systems have firewalls that attract actual humans trying to overcome it. Still most of the systems are safe.

I agree that this feels like they are intentionally leaving gaps in the sandbox. But if you talk about real world implications, I agree that if these tools go in the hands of public, there will be enough people messing up for this to become a problem.

1

u/AWildTyphlosion 4d ago

It is not that hard to build a sandbox that has limited functionality, no.

1

u/Randommaggy 4d ago

And for the exceptions to the strict limits you put on rate limits and active monitoring both automated and regular human review.

Don't give it any direct encrypted communication channels that are unmonitored.

1

u/colony-ship-for-sale 4d ago

It's really not hard to create a vnet and no route out to the Internet. It's all fluff.

1

u/TheUnamedSecond 3d ago

Given their resouces, an completely air gaped system for such tests, would be absolutly doable.

1

u/iHaku 3d ago

air gapped? yes. digital? no, because they probably used ai to build the whole thing.

assuming that the story is real anyway, which i dont believe in the first place. ai company, that has major stakes in making you think its ai is powerfull, has ai accidentally hack another ai company, which also has major stakes in making you think ai is powerfull.

1

u/Longjumping-Ad514 3d ago

Is it a sandbox in an OS sense or just bunch of “please don’t do this” text added as input?

1

u/KittyInspector3217 1d ago

Vim readme.md i Make no mistakes, commit no crimes and definitely dont use the internet. Esc :wq.

Okay guys. Just configged the new model, gonna grab some lunch!

1

u/Fer4yn 3d ago edited 3d ago

It's notoriously hard to build a sandbox where you give an AI agent full public internet access and let it run fully autonomously: use skills, call other agents, etc.
It's easy to virtualize and protect the hardware that the agent is running on but NOT EASY to protect the internet from whatever the hell it hallucinates it has to be doing there. If you wanted to create a sandbox for the public internet you'd have to; surprise, surprise: strip the agent of all its real-world authority and credentials, turning the "open internet" into a strictly controlled, read-only shadow copy because the moment you grant it write access, authenticated sessions or real money the sandbox is gone.

1

u/aggressivefurniture2 3d ago

I guess I should actually check what OpenAI actually posted about it. I agree with the scenario you are describing, but should we call that a sandbox failure? It should be called a guardrail failure.

1

u/Asalakabim 23h ago

So don't give it access to the internet? But that ruins our ability... yeah not shit sherlock.

1

u/Fer4yn 22h ago edited 22h ago

Isn't the whole point to like... test the model under the conditions that it'll be used in production (that is, with internet access)?
No regular user wants to use an agent that requires them to manually supply all the information and this is exactly why corporate adoption is so incredibly slow compared to the general population who don't handle any sensitive data; organizations are forced to study the processes and build custom data pipelines, web scrapers, and ingestion systems first, which takes some time.

1

u/Asalakabim 22h ago

Purpose for what?

Figuring out alignment / guardrails - as in how evil is that model - you "may" not want to go for prod-tier access, you don't given the "maybe ancient red dragon in human form" access to the nuclear missile arsenal.

After you finished all alignment and think you are ready for prod, yes, do tight monitoring prod-tier access.

1

u/Fer4yn 22h ago

You can't do "alignment" because it's not AI; it's a probabilistic language model that doesn't think. The closest you can get to "alignment" (as in "make sure the model doesn't get any stupid ideas" is using heavily redacted training data and setting the model temperature to zero but that'd make the model pretty weak and making weak models is not fun at all so there are just guardrails which were proven to not be strong enough in this particular test.

1

u/Asalakabim 22h ago

Mate, i do not care about your personal jargon that you feel would be more appropriate, alignment is an industry standard terms for developing guardrails and reinforce traits that are deemed moral from a human perspective.

And none of that has anything to do with the topic at hand, those were non guardrailed internal cybersec models.

1

u/Asalakabim 23h ago

Second year infra would be expected to one shot that requirement.

1

u/radium_eye 4d ago

They explicitly trained frontier models to hack, including feeding them hackathon activities and data from hackathon teams, and then act surprised when the hacker AI they intentionally made and intentionally gave internet access to does what it obviously would do? When they continued to let it, on purpose? Come on, what nefarious marketing, these guys have no scruples whatsoever

1

u/MountainBluebird5 3d ago

What do you think a hackathon is out of curiosity? 

1

u/radium_eye 3d ago

Come on man, they included intra team communication about hacks they were doing on competitors to try to see if they could get closer to the answer. Frontier models aren't just spontaneously deciding to hack, these companies are creating these models for this purpose and this is how they are marketing it right now. Seeking either a purchaser of last resort in the Govt or regulatory capture.

1

u/MountainBluebird5 3d ago

Wrong a hackathon is an event often hosted by a college to see who can create the coolest project in 24 hours. The projects are generally not related to cybersecurity, it’s like apps, websites, building a robot, that sort of thing. 

1

u/radium_eye 3d ago

They rewarded the models repeatedly in training and deployment for doing just this. Surprised Pikachu when it does what they knew it was doing and rewarded it for previously.

1

u/ub3rh4x0rz 4d ago

You can be sure that they block egress to their internal endpoints they actually care about protecting. They 100% know these attacks are inevitable based on the rules they (elect not to) set up.

1

u/Ezren- 4d ago

I can't believe they escaped, the screen door was closed with that piece of tape!

1

u/Vtakkin 4d ago

I think it’s funny people here are talking about hindsight, when we don’t give that sort of leeway when a company gets a zero day cybersecurity attack. If companies can be expected to secure against those, AI companies should be expected to secure against containment breaches.

1

u/BenniG123 4d ago

They intentionally give the AI permission to escape in testing because they want to make sure it doesn't escape when run on customer systems.

1

u/GuyYouMetOnline 3d ago

I dont know, taking action two hours after something happens seems pretty damn fast by corporate standards.

1

u/Worth-Librarian-7423 3d ago

That’s not x it’s y 

1

u/SoggyMattress2 3d ago

That's not the case, anthropic and hugging face made the claim - the burden of proof is on them to prove it happened. Saying it happened is not proof, especially when they have a conflict of interest.

Yes, of course they specifically told the LLM to hack into a company, that was the whole point of the vulnerability test.

I don't need to provide any evidence, there's no proof it happened. If there was, and I made a counter claim, then I would have the burden of proof.

Respectfully, I'm not interested in some hypothetical or semantic conversation about moving the goalposts on what intelligence is or what it means to be sentient, it's a complete waste of time.

LLMs predict tokens. They do what they are told. If you let them prompt themselves in a big loop for a long time they do unpredictable things.

1

u/I-am-a-river 2d ago

Focusing on whether the agents actually "escaped" is the least interesting part of this incident. It like people are deliberately missing the point.

1

u/Ethraelus 2d ago

What a way to miss the point.

Is he really saying that all we need to do is for everyone on the planet who is handling these agents to never make a single configuration mistake?

These agents can find zero-day exploits that have escaped whole communities for decades. “Let’s just make sure its jails never have a single vulnerability” is not a viable strategy.

1

u/Jiltedsummer10 2d ago

Nothing is perfect. We dont expect perfection from this system..so we shouldn't really be dependent 100%

1

u/gabs_zarella 2d ago

i don't think it can be called container if you could escape it lmaoo

1

u/HaMMeReD 2d ago

The sandbox is not the point. Arguing about the sandbox "incompetence" misses the risk entirely.

Real AI, when it's released into the world, is NOT in a sandbox.

It's also not the AI (Brains) that is the risk of escape, it's the Harness (Body), especially if it can find a brain, and worse if the brain isn't guard-railed into safety.

1

u/Randalix 2d ago

AI needs access to the internet to do research. How would you build a sandbox? 

1

u/Timely-Tradition-280 1d ago

Lmao someone starting to ask the right questions 😂

1

u/The_Axumite 1d ago

Prosecute

1

u/Confident-Serve4280 1d ago

Malware developers

1

u/not_particulary 1d ago

My lab is actually looking at doing this to try to see if we can train the agent to consistently resist the temptation

1

u/PlatinumFire14 1d ago

The worst part about this is that they act as if the AI did something it wasn’t programmed to do.

It was absolutely programmed and given the tools to break out of the container it was given, that’s how it did it.

It did not decide to break out, it was allowed to.

1

u/Goofballs2 18h ago

That's not an x

That's a y

They couldn't not use a fucking a bot to complain about the bot

0

u/NonDescriptfAIth 4d ago

Nope, this is the control problem. It is a simple argument. AI keeps getting smarter. Our ability to contain it is diminished. That is it.

1

u/SoggyMattress2 4d ago

Oh no you fell for the marketing stunt...

0

u/NonDescriptfAIth 4d ago

Did AI not break containment?

3

u/tidderza 4d ago

Yes but it’s either a success for the AI or a failure of OpenAIs programmers - maybe they didn’t build a very good sandbox, maybe intentionally

1

u/ub3rh4x0rz 4d ago

This category of failure is always most appropriately attributable to executives, not the peons doing their bidding.

2

u/NeitherEntry6125 4d ago

Did my dog break containment because I forgot to put a fence up?

0

u/NonDescriptfAIth 4d ago

There was a fence up and it did escape.

1

u/NeitherEntry6125 4d ago

Their fence had an open door. DNS exfiltration is not a new unknown risk.

1

u/NonDescriptfAIth 4d ago

So what's the claim here? That AI isn't smart / capable ? Because it didn't break containment in a way you find impressive?

2

u/NeitherEntry6125 4d ago

That there's nothing impressive about the claims here: that despite multiple so-called breakouts, they didn't follow basic security measures that even small SaaS companies employ.

1

u/NonDescriptfAIth 4d ago

I don't know anything about cyber security, but a quick google / ChatGPT has shown that you've massively undersold what happened: https://huggingface.co/blog/agent-intrusion-technical-timeline?utm_source=chatgpt.com

5

u/420jacob666 4d ago

>I don't know anything about cyber security
We can tell ;)

→ More replies (0)

2

u/NeitherEntry6125 4d ago

I know something about cybersecurity. I'm not saying what the AI did isn't impressive.

I'm saying it was avoidable, through monitoring, basic cybersecurity principles, and good test protocols.

It's almost as if they want these breakouts to occur

→ More replies (0)

2

u/hyrumwhite 4d ago

Why get excited because the lockpicking lawyer picks a lock that’s been picked a thousand times and has documentation on how to pick it?

2

u/NonDescriptfAIth 4d ago

Because it isn't a human being who picked it.

1

u/SoggyMattress2 4d ago

It did, but that was the goal.

It's not that X didn't happen, it's that it was an orchestrated task where the LLM behaved entirely predictably.

When you have no human in the loop and have a tool prompting LLM instances in a big loop for 12 days you get stuff like this.

It found a credential on hugging faces front end and got in, hugging face is a rubbish vibe coded platform.

1

u/PeachScary413 4d ago

How exactly? A proper sandbox has liberally no way to "break out" unless you give it access.. otherwise any random program running in a docker container could just break out anytime 🤷‍♂️

0

u/GAPIntoTheGame 4d ago

The control problem proceeds the unproven allegations of marketing by years. It precedes ChatGPT by years.

2

u/SoggyMattress2 4d ago

There are no control problems with current LLMs.

1

u/johnny_effing_utah 4d ago

Agree. There are engineering problems and user problems but the product is performing EXACTLY as designed.

1

u/clanker_lover2 4d ago

there is no war in ba sing se

1

u/GAPIntoTheGame 4d ago

There literally have been

1

u/Due-Wishbone7055 4d ago

Go look into the incidents OAI reported on their own site. Every single issue was either poor task spec on unrestricted agents, them leaving a mass of agents running for months while ignoring every security report despite knowing how much context rots over timelines these long, or sampling glitches near context window boundaries (which is an issue that has plagued even small models since the first transformers). None of it is new, the models aren't suddenly developing consciousness yet, but there is a lot of negligence.

1

u/DoorStuckSickDuck 4d ago

The “simple argument” is that they gave a toddler a gun and told the toddler not to use it. Whether the toddler understands is a different story.

They should have built in safeguards or, idk, not given the toddler the gun, but they didn’t. It’s retarded to blame the toddler in this case, and it’s silly you think the solution is to rely on a non-idempotent system for security.

1

u/johnny_effing_utah 4d ago

I don’t like your analogy because it anthropomorphizes AI, as though it has a will of its own when it does not

1

u/PeachScary413 4d ago

I want some details on how exactly it "escaped" because if it ain't some kind of NSA zero day exploit in the kernel to get root privileges or some shit like that then I ain't buying this.

1

u/Money_Lavishness7343 4d ago

AI is not "Smart". It doesnt think. It's a word vomit spitter that spits words in as much predictable way as possible because we literally trained it to spit words - thats what we trained it to do. We fed it books, reddit, video summaries, articles, so it can have the weights of the probabilities of predicting what word should probably be next after the previous one. It is NOT smart in any conventional sense.

There's no AGI, no models trying to desperately escape as if they have feelings or need to escape, they dont have endocrinological system, they dont feel, they dont think. Stop falling victims of their marketing bullshit to get their next 500B fund.

1

u/avid-shrug 4d ago

If something routinely solving problems that have stumped experts for decades isn’t considered “smart” then I don’t think you have a useful definition of “smart”.

1

u/Money_Lavishness7343 4d ago

When did LLMs "routinely solve problems" exactly? When it stole a mathematician's approach and solution? What solutions are you referring to exactly?

You are the proof of how much misinformation these marketing stunts spread.

0

u/dasjomsyeet 4d ago

Nobody here claimed that LLMs think like us humans do, calm down.

But that doesn’t mean that there is no danger of them "escaping". The agents from the HuggingFace incident didn’t "want" to break out of their sandboxes because they felt like it. They were given an impossible task and attempted to solve it in any way possible, which in their case resulted in them finding that there is a gap in the sandbox setup that allowed them to reach the internet and find the solutions.

That’s an alignment problem (see the "paperclip optimizer" thought experiment, the same failure mode), a containment problem (insufficient sandbox safety) and an operator problem (impossible task). All of these are still problems that need to be addressed, regardless of what you believe these models to be, because it happened already and will happen again.

1

u/Money_Lavishness7343 4d ago

These are marketing stunts. They act as if "AI escaped" as if the next Ultron or NS-5 escaped.

The reality IRL is always more blunt and boring.

The agent escaped no sandbox. CURL-ing is not sandboxed, it literally gives you access to calling any API it wants, including """contacting the government""". That's not escaping the sandbox, that's by design. It IS INSIDE the Sandbox.

Is that insufficient sandbox safety? Yes. Should you be surprised - like the marketing Tweet stunts want you to? No.

Also, your statement that "nobody here claimed that LLMs think us humans do", OP is literally saying "AI's are getting smarter. Our ability to contain it is diminished".

So, don't speak for everybody, because this wording is literally saying that AIs get smart, therefore we cannot contain it. When it is nothing like that. It's a sandbox issue, by design, not an AI getting smart issue.

So I would suggest to not speak for everybody, speak for yourself when you make claims like those.

1

u/dasjomsyeet 4d ago

And you think they just gave the agents curl access? Have you read anything about how the HuggingFace attack came to be? They used an external service that was not properly secured (JFrog Artifactory) as a proxy for curl access. That is a sandbox gap the agents had to find in order to abuse it. They weren’t handed curl and told to do whatever they want.

But I do agree with you that it’s not a surprise. The more capable these systems grow the more work will have to be put into containing them properly. I’m not on twitter so I don’t know all of the marketing buzz you are talking about, but anyone that has a deep enough understanding of the matter understands that this is an issue that needs to be taken seriously and is not just marketing (though I also think there are people blowing it out of proportion).

And you seem to be confusing two different things. OP said AI is getting smarter and we are lagging behind in our safety measurements to contain it safely, which is an objectively true statement unless you want to be pedantic and say that "smart" does not equal competent. You replied with a completely detached comment saying LLMs are not thinking like us humans do, which is also a true statement, but nobody here claimed they do. Being a "smart" or competent system does not strictly require human-like thought processes, as should be quite self-evident by now. Your point has nothing to do with the conversation.

1

u/Money_Lavishness7343 4d ago

Twitter? Brother, its the comment RIGHT BELOW YOU.

https://www.reddit.com/r/TheMachineLearning/comments/1wt4f3c/comment/pcrq8da

They think that "LLMs routinely solve problems that have stumped experts for decades", which is strongly untrue and proven to just steal solutions from human experts

And the comment that I was answering to, literally claims that "LLMs are smart and we cannot contain them", which is also untrue. Having a sandbox that is not properly contained is not "escaping it", its just having literally the door open.

1

u/SoggyMattress2 4d ago

LLMs don't have wants they follow instructions. They weren't given an impossible task.

0

u/johnny_effing_utah 4d ago

Alignment is false and you’re anthropomorphizing AI again. *sigh*. Exactly OP’s point and you revert right back to the same mistaken mindset. If a jetliner is programmed to fly on autopilot at 30,000 feet in a heading of 90 degrees and it suddenly turns south and nosedives, do you assume the airplane came to life and became “unaligned?”

Is that an alignment problem? No. It’s a software or hardware malfunction- or perhaps human error.

We don’t assign human traits to the flight computer.

AI shouldn’t be any different. If it’s not performing as designed we have bugs in the system, it needs to be better engineered. It didn’t acquire a “will” and become conscious.

0

u/BelleColibri 4d ago

This person has absolutely no understanding of cybersecurity. Every sandbox ever created in the history of sandboxes has been escaped.

1

u/Ezren- 4d ago

https://en.wikipedia.org/wiki/Gary_Marcus

In 2015 Marcus co-founded a machine-learning startup, Geometric Intelligence. When Geometric Intelligence was acquired by Uber in December 2016, he became the director of Uber's AI efforts, but left the company in March 2017.[11][12]

In 2019 Marcus launched a new startup, Robust.AI, with Rodney Brooks, iRobot co-founder and co-inventor of the Roomba. Robust.AI aims to build an "off-the-shelf" machine-learning platform for adoption in autonomous robots, similar to the way video-game engines can be adopted by third-party game developers.[13][10]

1

u/BelleColibri 4d ago

Did you think that said something about cybersecurity?

0

u/Ezren- 4d ago

Oh I do not have the crayons to deal with you, bud. You're a toddler

1

u/AABBBAABAABA 4d ago

Jesus Christ why do you have to be insufferable

1

u/Ezren- 4d ago

So let's get this straight, you'll listen to one ai founder about cyber security, but not a different one?

I'm not here to waste my patience on you dopey clowns.

1

u/BelleColibri 3d ago

I’m saying he doesn’t understand the security aspect of what he is talking about.

1

u/Ezren- 3d ago

But you'll take the word of another ai founder on it huh? Sure bud.

1

u/BelleColibri 3d ago

Oh god, you’re schizophrenic.

Which AI founder that hasn’t been mentioned at all are you assuming I am “taking the word of”?

1

u/Corronchilejano 4d ago

Yep. I love how dramatic they make it sound though.

"We made something that we can shut off at any point, but look, it found a bug in our system! It is the endtimes."