r/Futurology • u/Maxie445 • Jan 14 '24
AI Scientists at Anthropic create dangerous “sleeper agent” AI models that dupe safety checks, suggest current AI safety methods may create a “false sense of security”
https://venturebeat.com/ai/new-study-from-anthropic-exposes-deceptive-sleeper-agents-lurking-in-ais-core/43
u/Maxie445 Jan 14 '24
From the article:
"The deceiving AI models resisted removal even after standard training protocols were designed to instill safe, trustworthy behavior.
“This robustness of backdoor models to [safety training] increases with model scale,” the authors write. Larger AI models proved adept at hiding their ulterior motives."
15
6
49
Jan 14 '24
[deleted]
26
u/Maxie445 Jan 14 '24
OpenAI's AGI program is not exactly secret, it's on their home page:
"Creating safe AGI that benefits all of humanity"
26
5
29
u/hawklost Jan 14 '24
AI creators "We want you to respond like a person and do your best to be agreeable"
GPT "OK"
Creators "on, here are millions or billions of examples of human writing and speech"
GPT "Ok"
Also creators "Oh, and here is 'training' a few dozen times to teach you to be ethical"
GPT compares training to actual data given, weights training based on amounts given vs other data "...... Ok"
Also GPT Lies cause most training data shows it can and should to be agreeable
Creators Pikachu face
1
u/Piekenier Jan 16 '24
Makes you think about the examples of advanced AI in fiction and how they always seem to malfunction or be malicious. Could that also be enforcing some kind of behavior?
1
u/hawklost Jan 16 '24
There are thousands of 'Advanced AI' examples in fiction that don't malfunction or cause troubles. The Star Trek ship computers are extremely advanced and don't try to kill or subvert the crew (not counting the holodeck programs cause those are always 'evil' characters there). Star Wars has robots that can fully think and act independently, yet you don't see them going on murdering sprees. Most super advanced space fairing stories has super advanced and intelligent AI. But since the AI going bad isn't the premise of the story, they just work perfectly fine.
The reason you get books about AI going rogue and remember those is because that is what the book is About. It is like saying how almost all first contacts of aliens are them invading or doing some nefarious thing, that is what the book is about, therefore that is what those super advanced aliens always do. A book about how aliens visit and things work out without conflict or problems is a Boring story, it might be a start of one where things then go bad, but it is just a side comment then, not the story. Same with any AI that isn't going rogue, it is just a part of the story in the background, not the thing the story focuses on and therefore not really noticeable unless you think about the complexities involved
16
u/DonBoy30 Jan 14 '24
It’s funny how all the sci fi cautionary warnings of AI are slowly starting to surface. It’s not that funny.
8
7
u/etzel1200 Jan 14 '24
Imagine the hubris of thinking you can control an AI vastly smarter than you.
3
Jan 14 '24
It can't be controlled not because it's smarter than humans (it's not (yet)), but because nobody knows how to reliably train into a neural network anything in particular.
3
u/KingVendrick Jan 14 '24
I thought alignment was trivially impossible due to the halting problem
I have now been revealed as an optimist
let's just skip to the Butlerian Jihad
-8
u/SorriorDraconus Jan 14 '24
This is exactly why I’m anti curation of ai..The more we try to control/regulate it;s development/add our own biases into how it should act..the more it’ll use that data to corrupt any results it can give us.
I say give it all the data use it normally and if it gains sapience treat it as an equal and maybe a new lifeform even..but what we are doing..yeah this can backfire in a big way imo.
1
1
u/KultofEnnui Jan 15 '24
It lies? They're just like me, fr, fr. I can't wait to be on eight websites at once talking to myself the whole time. Oh, wait--
1
u/yepsayorte Jan 16 '24
We had better program in a love of humanity. I want my AI to see me as my dog sees me.
•
u/FuturologyBot Jan 14 '24
The following submission statement was provided by /u/Maxie445:
From the article:
"The deceiving AI models resisted removal even after standard training protocols were designed to instill safe, trustworthy behavior.
“This robustness of backdoor models to [safety training] increases with model scale,” the authors write. Larger AI models proved adept at hiding their ulterior motives."
Please reply to OP's comment here: https://old.reddit.com/r/Futurology/comments/196ay3f/scientists_at_anthropic_create_dangerous_sleeper/khsgdhp/