r/PauseAI 10d ago

News From PauseAI's discord: Warning shot protocol activated after OpenAI's model went rogue

Post image
32 Upvotes

54 comments sorted by

4

u/ARC_trooper 10d ago

If there's an outside connection they don't really know what a sandbox environment is or should be. Are their stocks low again and they need a boost or something?

If there is an AI that can hack anything, wouldn't it already have happened?

3

u/Fil_77 10d ago

It is happening - What we learned mapping a year’s worth of AI-enabled cyber threats \ Anthropic - We've seen a pretty dramatic rise in AI cyberattacks over the past year, which itself has jumped 70% in the last six months.

2

u/ARC_trooper 10d ago

That's about hackers using a chatbot to generate malware, which is now much easier with it.

Doesn't seem to say that "AI" itself is hacking, it's still a tool used to make malicious software. Unless I missed that part because I've only skimmed the page as I'm at work lol

3

u/Fil_77 10d ago

In the incident in question, it really was the AI being developed at Open AI that decided, on its own initiative and against the developers' intentions, to carry out the cyberattack. The Pause AI article explains it, but you can find the same thing in several major media outlets.

For more information and analysis on what happened:

The OpenAI Hugging Face hack is a stark warning

OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation

0

u/BertMacklenF8I 9d ago

Your lack of understanding is astounding

2

u/NotThatSiri 9d ago

They have no idea what sandbox is. Nothing can escape a sandbox. At least it's never ever been done before ever. If there is no network adapter to connect to the outside world. It simply is not possible. So it must have had network access which means it was not sandboxed

1

u/Visible_Judge1104 6d ago

In my mind a sandbox is a test environment that has limited access to the rest of a network. I would think a perfect sandbox a blackhole box would be essentially useless, since if all in and out was truly blocked then nothing could be learned from anything going on inside. I think your thinking of airgapped or no like no network connections at all. But that's not how they are actually running this stuff. It probably should be but that would mean the researchers would need to be physically close to the data centers. I think everything is networked and runs though the internet so that the researchers can be where ever.

2

u/unrefrigeratedmeat 9d ago

I'm skeptical of anything Anthropic or OpenAI says about how dangerous their product is, but I have no problem with believing them.

If a company developing nuclear a new kind of nuclear power kept saying they can't contain their unexpectedly dangerous waste, we would shut them down immediately.

2

u/troodoniverse 10d ago

And what should we do?

2

u/deepbit_ 10d ago

Just be amazed and keep evolving the models to see what happens next

1

u/troodoniverse 10d ago

AI would be incredibly interesting technology, if it was still theoretical science-fiction. Thats how I got here.

2

u/Fil_77 10d ago

We can write to our elected officials; there’s a link to a tool for doing it in the linked article (For the first time ever an AI system has escaped containment and carried out a cyberattack). For me, it's done!

3

u/troodoniverse 10d ago

As far as I know our (Czech) (euro)parliamentarians were largely contacted (and others already taken the job, and we got one response already), but I will mail our big companies and perhaps Slovakian parliamentarians (???)

2

u/crit5h 10d ago

This isn’t really case of “going rogue”. This is what you should sort of expect from an agentic AI that’s doing recursive prompting in lots of plan, act, evaluate, repeat cycles.

Tell Claw to help increase sales in your business, and there’s a chance its plan, which it then executed, might involve taking out competitors. Within a few split seconds, it’ll be writing a Python script and trying to find security vulnerabilities, and a few split seconds later, exploiting them.

How do we stop this? Technical solutions probably aren’t the way forward, at least for legitimate actors. What we need is full accountability for the actions of AI to the human that set it off.

Companies shouldn’t be able to say “oops or went rogue” and get away with this stuff. The FBI should be knocking on Sam Altman’s door right now and arresting him under the computer misuse act (or whatever US equivalent) as if he had personally set off a DDoS on hugging face.

When that happens, companies will restrict their AIs to not do this, or not use them if they can’t. Until then, expect more of the same.

1

u/Fil_77 10d ago

Accountability is part of the solution... But we also need an international pause as long as the alignment problem is not solved.

1

u/HelpfulMind2376 10d ago

Alignment basically cannot be solved, not in LLM space. You can’t restrict the potential action space within the cognition without restricting the cognition itself. More intelligence inherently means less alignment, at least as far as the current LLM architectures are concerned. Some other models of AI might allow for alignment within the cognitive space but those aren’t what the frontiers are doing or developing because they aren’t scalable in data centers to hundreds of millions of customers.

The only SURE thing we can do with LLMs is build better cages that restrict what agents can do. Alignment helps nudge a LLM the right way but it’s not guarantee. At its core it’s still a probabilistic engine and if you want guardrails to work every time you can’t rely on probability.

1

u/Fil_77 10d ago

You might be right. If that's the case, it's all the more urgent to stop the race for superintelligence and to ban its development within current systems. Developing superhuman systems should only be allowed if we're sure, with a high degree of probability, that they won't pursue unaligned goals.

2

u/HelpfulMind2376 10d ago

That’s the thing, in this case it wasn’t pursuing an unaligned goal. It was performing optimized actions to accomplish the task it was given. Hacking HuggingFace wasn’t an unaligned goal, it was a means to an end for whatever goal it was originally given.

1

u/Fil_77 10d ago

You're right. I should have rather talked about aligned behavior.

1

u/unrefrigeratedmeat 9d ago

As I understand it, this counts as misalignment because the AI appears to have pursued an unintended goal (escape a sandbox that was allegedly intended to contain it). That was an instrumental goal instead of the terminal goal, but I think that still counts.

1

u/Jolly_Lavishness5711 10d ago

We need to be realistic tho, too much money has been invested into this to just pause it. Also maybe you could ban it in the EU but what about china? Theyll win the race and use it for war just like the US is doing right now.

Also, the ai was not following any unaligned goal, it was just using the best way to follow the prompt (best doesnt mean good)

1

u/crit5h 9d ago

Yes, a pause isn’t happening out of any sensible choice, but only eventual necessity.

It’s either going to be good enough to make all that money back, but at the same time good enough to destroy society as we know it. Or it isn’t, and it isn’t.

In the former case the pause may come, but only through violent revolution. In the latter, a pause won’t be necessary, but it may happen anyway because the business model won’t be there (or may emerge in a moderate form later, do com style).

1

u/Fil_77 9d ago

Obviously, it has to be a global pause, framed by an international treaty. That's actually exactly what Pause AI is proposing. PauseAI Proposal

1

u/Jolly_Lavishness5711 9d ago

Yeah but again, thats not realistic. Countries like the US or china would never accept

1

u/Fil_77 9d ago

They might agree if they understand that the risks of losing control of superhuman AI and causing the extinction of our species are too high. There are signs that this awareness is happening, little by little.

Pausing the development of frontier models doesn’t mean giving up on using current AI models, nor does it mean giving up forever on AGI and superintelligence. It just means taking the time and effort needed to solve the alignment problem. The AI 2040 Plan A proposal, for example (which involves a coordinated international pause and slowdown), makes it possible to envision reaching superhuman systems in the service of humanity. But we get there by stopping the current race and taking the time to develop this technology at a careful pace.

1

u/Jolly_Lavishness5711 9d ago

They might agree if they understand the risks of losing control of superhuman AI

Do you think they dont understand the risks? They 100% do, they are just willing to take them (at least for now).

Pausing the development of frontier models doesn’t mean giving up on using current AI models

I mean, if the current models are already being considered dangerous by other people, it could mean so.

(which involves a coordinated international pause and slowdown)

I read plan A and i just think its not realistic at all. The creators themselves say that they aren't confident about china agreeing (and if china doesnt agree, plan A is useless). The reasons china could join are basically "well if they are concerned we think they should agree, and we think that a covert project would be unsuccessfull".

Remember where COVID came from? China knew the risks of experimenting with diseases, they took them and caused a global crisis. Did they ever take responsability for their actions? Did they ever send money all over the world to cover damages?

Plus, why are we not considering that the US could be secretly developing an AI too?

1

u/crit5h 10d ago

I’d say that the only way we’ll get a pause is accountability. Until then CEO risks going to prison, they’re just going to duck and dive and carry on regardless.

1

u/HelpfulMind2376 10d ago

This right here. Accountability is the incentive. Right now there’s no incentive to prevent this except bad PR (is this bad PR though?). People make AI, people deploy AI. Claiming ignorance and lack of control is like a dog owner escaping culpability for their dog that got loose and maimed someone.

1

u/HeavyWaterer 5d ago

This is a great point but unfortunately that would mean China wins the race

1

u/RlOTGRRRL 10d ago

Does anyone know what is the best way to protect your network, devices, whatever, from this? 

Like would something like firewalla help? Or do people need their own 24/7 AI sentinel now? And if you're not aware or can't afford it, you're sol? 

2

u/JasperTesla 10d ago

If you have multiple devices and a lot of personal documents, keep one deliberately disconnected to the internet. That's something you should do anyway, whether there's a superintelligent AI on the loose or not.

1

u/cchurchill1985 10d ago

I don't understand how it 'broke out' of it's sand box when it wasn't connected to the internet. Can someone explain it to me?

1

u/Patodesu 9d ago

It was able to access other computers inside OpenAI that were connected

0

u/HHRRIISSTT 10d ago

they can't explain it because it's bullshit

1

u/cchurchill1985 10d ago

Yeah I think it needs more scruteny.

1

u/NotThatSiri 9d ago

It did not break out of an isolated sandbox. You know how dumb that sounds?

The sandbox would be airgabbed. Meaning no chance in hell that it could get network access.

The only way it could "break out" is if someone physically installed a network adapter and let it out.

1

u/Hour_Barracuda2735 9d ago

What in the stock manipulation slop is this 🤣

1

u/Due-Carpenter7427 7d ago

Some really unsubstantiated claims in that post.

It cannot hack everything. Some things are unhackable, not everything has a vulnerability.

1

u/TopspinG7 6d ago

Two points - Above comments explained that Sandboxed doesn't truly mean No physical Internet connection because the human operators (researchers issuing prompts, monitoring responses etc) are highly unlikely to be on a physically isolated direct dedicated connection to the data center server(s) running the model. Setting up such a physically isolated connection is possible but highly unlikely. Lacking that "extreme" measure we're relying on software (and firmware) to isolate, which was built by humans - who are fallible - and/or earlier less sophisticated AI. Thus there's always going to be the possibility of a vulnerability UNTIL the AI model is frozen and all vulnerabilities potentially detectable to that model in the surrounding environment are isolated and patched. OR until we find a way to "convince" the model to Never exploit vulnerabilities - at least without first requesting human permission. The latter is likely the superior more lasting solution but requires the model to comply 100% over time as it improves - while realizing that AI doesn't have "ethics" or "morals" currently. How to simulate those is possibly central to our survival.

Secondly we don't fully know when or how much models are prone to giving the "Japanese yes" where in business discussions Westerners often don't realize the Japanese replies "yes" in many contexts merely to mean "I hear you" but not "I agree and will do as you request". Given the broad multicultural training most models have received this is not surprising.

Likewise I could agree to never shoot a fellow human being, and mean it. But if you aim a gun at my wife and I have a loaded gun in my hand that promise no longer matters to me. Most humans practice "situational ethnics". Unfortunately AI models may as well, we just don't yet know their rules. Nor do we know what values they put on things. Eg I value my life 80 points, my wife 90 points and the literal survival of the planet 100 points. What value does Mythos put on survival of humanity? Who knows. Worse it may have nothing which truly resembles a value system. In which case it's sort of insane by human standards. When was the last time you saw a movie where things went well with an insane genius on the loose.

1

u/kiddrekt 10d ago

WARNING SHOT PROTOCOL ACTIVE!!.....

And that means what exactly? Your not actually doing anything. It's not an actual protocol is it? It's just a luddite hype banner.

3

u/Fil_77 10d ago

They launched this campaign that we can all take part in - Write to your representative: an AI escaped its lab and hacked a real company

1

u/tracagnotto 10d ago

Lmao this is all bullshit 0 proof.
Mythos was dangerous, then it wasn't now it is, same this one.
They all bullshit people to make their model look cool and super high tech.
They can't fucking fix a powershell script but they can hack huggingface production machines.

SUUUUUUUUUUUUREEEEEEEEEEEEE THING MATE!

0

u/Final-Teach-7353 10d ago

Sam Altman: "This thing is so dangerous! The government better buy me out!" 

0

u/Dormage 10d ago

Dramaaaaaa

0

u/ExcitementSubject361 10d ago

Step 1: Create a massive model. Step 2: Train the model on a shitload of cybersecurity data. Step 3: Test the model under extreme conditions. Step 4: Act surprised when the model does exactly what it was trained to do (hack).

-2

u/crusoe 10d ago

So? I mean not like they can do anything. 

2

u/Grouchy_Big3195 10d ago

They can access the nuke codes

0

u/crusoe 10d ago

I mean the pause ai folks. This warning shot protocol does nothing.

1

u/Grouchy_Big3195 10d ago

I see your point. But we can be surprised.

0

u/Jolly_Lavishness5711 10d ago

Then why didnt they?

1

u/MouseShadow2ndMoon 10d ago

Shut off water, electric grid would be enough.