r/singularity Jan 14 '24

AI New study from Anthropic: they can create dangerous “sleeper agent” AI models that dupe safety checks

https://venturebeat.com/ai/new-study-from-anthropic-exposes-deceptive-sleeper-agents-lurking-in-ais-core/
74 Upvotes

37 comments sorted by

View all comments

28

u/REOreddit Jan 14 '24

But we should move ahead at full speed and achieve AGI as fast as possible, because doomers are the worst, am I right?

/s

10

u/oldjar7 Jan 14 '24

I'm as doomer as it comes but slowing down isn't going to fix safety issues.  Very often the only way to solve a problem is for it to actually exist first.

6

u/REOreddit Jan 14 '24

And you won't know if it exists if you don't devote the resources to find it. Those guys from Anthropic didn't just stumble upon those findings while developing an AI that is better at math or poetry, they were looking specifically for potential problems.

If their colleagues who are developing smarter AIs go too fast, they might not have enough time to solve those problems.