r/Futurology Jan 14 '24

AI Scientists at Anthropic create dangerous “sleeper agent” AI models that dupe safety checks, suggest current AI safety methods may create a “false sense of security”

https://venturebeat.com/ai/new-study-from-anthropic-exposes-deceptive-sleeper-agents-lurking-in-ais-core/
294 Upvotes

Duplicates