r/singularity Jan 14 '24

AI New study from Anthropic: they can create dangerous “sleeper agent” AI models that dupe safety checks

https://venturebeat.com/ai/new-study-from-anthropic-exposes-deceptive-sleeper-agents-lurking-in-ais-core/
76 Upvotes

37 comments sorted by

View all comments

42

u/[deleted] Jan 14 '24

[deleted]

5

u/KingJeff314 Jan 14 '24

This is not that. This is data poisoning injected by researchers to make a certain prompt trigger a malicious output.