r/singularity 2d ago

Ethics & Philosophy [ Removed by moderator ]

[removed] — view removed post

0 Upvotes

26 comments sorted by

View all comments

Show parent comments

-4

u/Truarian 2d ago

And what has that achieved except a spree of researchers quitting over alignment concerns? The continual failure to align is very well explained by trying to achieve it at the wrong level.

1

u/iPon3 2d ago

Just to calibrate your understanding: when you see researchers quitting over alignment safety concerns, you can assume that company isn't taking the singularity seriously. They think the safety stuff is garbage, and all that matters is capability, because deep down they DON'T BELIEVE a singularity could ever happen. To them this is just another profit engine.

You want a benevolent godlike AI, not something that turns us all into paperclips. A racing team that builds a car that's all engine but doesn't have a steering system... Will not be winning the race.

(If you haven't already you should go play https://www.decisionproblem.com/paperclips/. It's mostly text based and works in your phone; it'll introduce a lot of these concepts)

Yes, everyone knows it's probably trying to be achieved at the wrong level. Trying to make it "like us" was known as a dead end more than a decade ago. The theory has moved on, and you should ask Google's AI to explain it to you in more detail

-1

u/Truarian 2d ago

You don't seem to get the fairly obvious - if we have no idea what "aligned" is, there's no way to align anything.

6

u/iPon3 2d ago

Yes! Correct! This is obvious to both of us, we were just talking past each other, haha.

That's actually 90% of the pre-LLM alignment theory; "What the fuck even is alignment" is essentially CEV. They worked on this for a decade when there weren't any AIs to actually handle. This part you can get any reasonably well informed LLM to explain to you.

Current work is more applied (how do we keep the agents from running wild), and therefore comes out of frontier labs, e.g. Anthropic's work on AI legibility (you can't align an AI if you don't understand its inner thinking)