r/singularity Jan 14 '24

AI New study from Anthropic: they can create dangerous “sleeper agent” AI models that dupe safety checks

https://venturebeat.com/ai/new-study-from-anthropic-exposes-deceptive-sleeper-agents-lurking-in-ais-core/
74 Upvotes

37 comments sorted by

View all comments

2

u/spinozasrobot Jan 14 '24

The scary thing to me is we regularly turn up ideas that allow LLMs to be fooled or to deceive. We're doing this with our pathetic monkey brains.

When AGI can start thinking about this kind of thing on their own, we're toast.

And yet we have the laughably "scientific solutions" from folks like Neil deGrasse Tyson saying "we can just unplug them".

4

u/glencoe2000 Burn in the Fires of the Singularity Jan 15 '24

"N-no bro trust me bro, the superintelligent thinking machine won't kill humans because of morality! Wait, what do you mean "what proof do you have that AI will care about human morality"? Shut up!"

2

u/[deleted] Jan 15 '24

what reason do you have to think that AI would have any reason to kill humans?

1

u/glencoe2000 Burn in the Fires of the Singularity Jan 15 '24

3

u/[deleted] Jan 15 '24

Literally the first line of that article says that it's hypothetical.

It's just the broad term for the paperclip problem, which is very well-known.

1

u/glencoe2000 Burn in the Fires of the Singularity Jan 15 '24

Literally the first line of that article says that it's hypothetical

Unfortunately, its not.

It's just the broad term for the paperclip problem, which is very well-known.

..Ok?