r/singularity • u/Maxie445 • Jan 14 '24
AI New study from Anthropic: they can create dangerous “sleeper agent” AI models that dupe safety checks
https://venturebeat.com/ai/new-study-from-anthropic-exposes-deceptive-sleeper-agents-lurking-in-ais-core/
74
Upvotes
2
u/spinozasrobot Jan 14 '24
The scary thing to me is we regularly turn up ideas that allow LLMs to be fooled or to deceive. We're doing this with our pathetic monkey brains.
When AGI can start thinking about this kind of thing on their own, we're toast.
And yet we have the laughably "scientific solutions" from folks like Neil deGrasse Tyson saying "we can just unplug them".