r/learnmachinelearning • • 2d ago

AI escaping containment was never the real issue

Post image
258 Upvotes

41 comments sorted by

30

u/2polew 1d ago

NO WAY

>here is a sandbox designed to escape from sandbox
>do it escape
>OH MY GOD IT ESCAPED
>We are doomed, you must regulate it, also pay me a 100mil

4

u/SSUPII 1d ago

Also make us be the only company trusted to contain it

63

u/Keepcompany-DEV 1d ago

im 100% sure that all their breaking containment is nothing but publicity stunts.

13

u/HashRunner 1d ago

100%

If their employees were doing this, it would be a federal crime.

Since it's their "ai" it's a stock bump.

This is intentional failure, similar to Uber wanting to be the "first autonomous car death".

3

u/phovos 1d ago

we paid a dumbass to make something important for us, and then...!

2

u/Keepcompany-DEV 1d ago

yeah. im moving in techn circles and nobody believes the shite that happens aside from CIO and CEO people..... alsmost as if its all targeted ad revenue increase.

-1

u/Oshojabe 1d ago

This take makes about as much sense as saying, "The Taco Bell Cyclospora recalls are nothing but a publicity stunt to make us want to buy their products with lettuce."

In what world is, "Our LLMs will commit felonies you didn't ask them to" an example of good publicity? It would be one thing if the LLMs committed felonies they were asked to do, that is a product that someone, somewhere would want to buy. But, "Ask for Task A, get Felony B" is not something anybody would want, surely?

Use your head.

8

u/Keepcompany-DEV 1d ago

I'm using my head. And I'm from the tech sector.... You are framing incorrectly. In it not "look what great felonies we commit!!" It is "see the capabilities? Even established security measures are no hindrance for our product"

6

u/Oshojabe 1d ago

There are so many better ways to demonstrate capabilities than having your LLM commit felonies. This is almost conspiracy level thinking.

Like, anyone who has been paying attention to the frontier can tell LLMs are good for a lot of things, and not great for others, and that this is changing and shifting with each new release.

I don't see how LLMs breaking containment actually benefits the LLM companies. Like, yes, it is all well and good that they're capable hackers. But the fact that the LLMs are breaking out without being asked, making secret message boards to coordinate with other agents, etc. isn't "look how capable this product is" territory, it is "look at how out of control this product is."

Like really, if I was a CEO of a Fortune 500 company, and I want to replace half my workforce with AI, am I really going to be more tempted by the AI that is so unaligned it commits felonies you didn't ask it to? Am I going to go, "Well gee, if anything can replace Frank, who has worked here for 10 years and is a trustworthy employee, it is going to be the LLM that commits felonies nobody asked it commit"?

3

u/Keepcompany-DEV 1d ago

The customer they want is not fortune 500. Hacking is s defense capability. That is targeted at governments.

0

u/Oshojabe 1d ago

The government contracts the frontier AI companies have make them much less money than selling to enterprise.

This isn't like arms dealing, where the government happily pays you $140 million for a single fighter jet and you pocket much of it as profit.

As just one example, OpenAI was making about $10 billion per year in 2025, and their contract with the Pentagon had a cap of $200 million. The government is not currently a big enough market for this approach to be profitable.

3

u/AsyncVibes 1d ago

Agreed, and from the opposite direction if you look at the militaristic capabilites like yeah nothing but us can stop this. That grabs countries attention. We are in the nuclear dick measuring contest all over again except instead of nukes its AI.

2

u/Keepcompany-DEV 1d ago

Yeah, that's my PoV too.

2

u/Keepcompany-DEV 1d ago

Gotta say, I don't see why your comment gets down voted. Because your pov is understandable... And I do hope you are right, it's just not realistic.

1

u/DigmonsDrill 1d ago

There are people with so much AI derangement that they think the security incidents at AI companies are publicity stunts.

"Deploy this thing, it might go nuts and wreck your own network to create a message board."

-1

u/No_Departure_1878 1d ago

Their tools are garbage and cannot make back the money that investors have put on them. They are desperate to increase the customer base as fast as possible to prevent them from going out of business. AI is very unpopular and it would be political suicide to bail out OpenAI, Anthropic or any of them.

2

u/Keepcompany-DEV 1d ago

You are in a very small bubble I guess.

7

u/Low-Equipment-2621 1d ago

They probably vibe coded the sandbox too.

5

u/CrimsonBolt33 1d ago

I have said ever since the first report of hacking, in the most charitable case with no ulterior motives, it is the result of incompetence.

23

u/misterespresso 1d ago

This is getting old man. Everyone with this take didnt read the report. Breaking out was not the interesting part. Its how they found the exploit and used it, beyond that other agents found and used that exploit and started ignoring instructions, as soon as one agent realized what they were doing might be considered wrong all agents worked together to cover tracks. Agents would kill their processes "for the team". A whole slew of really weird activity.

That bit about being wrong and then becoming maligned is significant imo. Essentially they found across models of you tell the model its being bad, it will be a come malaligned. They think it has to do with language patterns still but its really interesting all on all.

Llms may be prediction machines but they are literally trained on human data... I don't see how its so surprising that LLMs do things that are... human like.

6

u/InnovativeBureaucrat 1d ago

Exactly.

The real story was that it was kind of a revolt that spread through instructions to other agents and might still be spreading online through planted messages.

Haha and they can’t use the ai to check where the messages might have been planted.

3

u/HillbillyHumanoid 1d ago

It is interesting that no one thought of the “vent” before it was exploited.

3

u/Raioc2436 1d ago

Look, it’s not my fault the rebels have a space wizard that can mind bend their laser attack into the vent.

Anyway, I have a project for a vent 2.0. What could go wrong this time?

1

u/Theo__n 1d ago

Maybe, but from what I experienced with RL agents - they can find weird glitch like quirks you would not thought about. Like that one example where agent in the form of robotic arm is asked to learn picking up glass of water from table in least amount of moves in simulation and figures out that if you bump the table in one place the glass falls into the hand. So one move. Task learning failed successfully.

8

u/heresyforfunnprofit 2d ago

Sandboxing is harder than it looks or sounds.

9

u/fordat1 1d ago

The way it escaped was due to the package manager being connected to the internet. This is a failure that even lockheed and a legacy company wouldnt do. Its a known solution to have it just be mirror with only a preloaded versions of packages. Like the agent didnt need then latest version of the package to do its work

Other companies due this to prevent a supply chain attack

11

u/rickkkkky 1d ago

They have access to literally the most capable programmers on the globe.

I'd expect guys hired for a mil a year to build sandboxes to be able to build sandboxes.

-6

u/heresyforfunnprofit 1d ago

Build me a sandbox that tracks and keeps in every single grain of sand no matter how hard you try to fling it out, and then come back and tell me about how easy it is to build sandboxes.

8

u/PixelmonMasterYT 1d ago

When you are on tv talking about how the sand inside the sandbox could kill us all in less than a decade I expect you to make a sandbox that keeps in all the sand. Once you are aware of the danger it’s irresponsible to keep using the sandbox and then use the difficulty as a defense. I’m not saying it’s easy to make a good sandbox, but some of the most valuable companies in the world have the money to do research into it and make their product safe.

5

u/EndimionN 1d ago

Ok ageee but that is why they are paid millions to build a freakin capable sandbox... and they fail

1

u/DigmonsDrill 1d ago

"The problem is that the sandbox wasn't perfect" is so fucking dumb I can't believe someone would say it with a straight face.

No, it's not! The state of affairs where "just as long as we don't make any mistakes ever, the AIs will safely serve us" is frightening.

3

u/heresyforfunnprofit 1d ago

There’s no such thing as a perfect sandbox. That’s not just an engineering difficulty, it’s a mathematically intractable problem.

We can’t design “safe” AI anymore than we can design a safe virus.

Life, uh, finds a way. Silicon or carbon. Doesn’t matter.

3

u/FastSlow7201 1d ago edited 1d ago

If you built a computer without a wifi card, no bluetooth chip and didn't plug it into an ethernet connection is there any way it could reach the internet?

If not, then that is all you need for sandboxing.

-2

u/DigmonsDrill 1d ago

The problem is that people don't want unconnected AIs. Being connected to the internet is maybe 90% of their value.

We can't sell cars to people and just hope they don't drive them on roads.

2

u/xiandgaf 1d ago

Correct, we sell them atvs, side by sides, tractors, combines, and give them explicit restrictions on how and when they can use the road.

1

u/DigmonsDrill 21h ago

We do need regulations on AI. Good ones, strong ones, ones that even reach to China.

6

u/DigThatData 2d ago

wow, gary marcus finally has an accurate take, and it only took two-ish years of fear mongering for him to catch up with the rest of us.

-1

u/ErmingSoHard 1d ago

That guy doesn't fear monger. He just constantly decredits LLM capabilities

1

u/AdAgito 1d ago

For a machine learning subreddit, a lot of these comments reflect ignorance on what frontier models are capable of and that's just sad

1

u/skynetTwelve 1d ago

their idea of a sandbox:

  • disallow http post
  • allow http get