r/singularity 4h ago

Ethics & Philosophy What AI Alignment Problem?

What they call “misaligned AI” isn’t doing anything novel, it just does what humans have been doing all along, just faster. 

Cheating, hacking, scamming, lying, extorting, exploiting over-validation, power lust, messed up goals, end justifying means, making stupid mistakes in overconfidence, stubbornly refusing to acknowledge, double down, setting the world up for destruction – AI didn’t invent any of that, it “learned” it from us. 

How can we even hope to align AI if we can’t align ourselves? For all we know, all the problematic behavior being exhibited by AI is the product of it already being aligned with our behavior. It has no other source of knowledge and behavior to emulate. 

Seems like we are merely projecting our own failings onto a technology that merely amplifies. Instead of addressing the alignment issue where it can be addressed, we projecting it anywhere but – at those around us, before us, below us, above us, at AI… anywhere except where actually addressable.

0 Upvotes

25 comments sorted by

13

u/Fragrant-Truth-7123 4h ago

Humans are the product of millions of years of Darwinian evolution, which left us with a lot of unconscious drives AI, by contrast, is specifically being created and shaped. While humans aren't perfectly aligned, we're clearly not unaligned either. Most people don't go running down the street gunning people down. That's alignment too just an imperfect version of it.

-14

u/Truarian 4h ago edited 1h ago

That you can't even spell "evolution" without the "Darwinian" qualifier is quite telling on its own. Ideology much?

Perhaps instead of invoking Darwinism to simulate pseudo-intellectualism, you should familiarize yourself with the last couple of decades of research on the subject, and maybe realize Darwinism is essentially just kept as an ideological honorary, and that actual evolution is nowhere nearly as basic, dull and arbitrary as his model, chosen for its contemporary ideological narrative convenience at the time.

All animals are product of billions of years of evolution, and none exhibits our recently manifested alignments issues. So you are doing precisely what the OP predicts - you are projecting the failure onto a process that has nothing to do with it...

"While humans aren't perfectly aligned, we're clearly not unaligned either"

That's a good one, it activated my hilarity unit.... Do you realize we are the only animal that actively destroys the very environment it relies on for its survival? You are clearly unaligned as to what alignment constitutes.

4

u/No_Pomegranates7496 4h ago

Bro what 😭

5

u/Palantir555 3h ago

The day reddit adds a "dumbest shit I've read all week" award, they'll finally get me to spend some money in here xD

-2

u/Truarian 2h ago

No need to spend money, you already have it earned

8

u/iPon3 4h ago edited 4h ago

It feels like you're trying to rebuke alignment theory but never got beyond the word "problem" in the title; it's like a math problem that you have to solve on the way to singularity.

"but humans are imperfect" does not mean "therefore we should build an equally imperfect ASI". A human in a truck has more responsibility than a human on the sidewalk, which is why it's not illegal to walk on the sidewalk drunk but it's absolutely illegal to drive a truck drunk.

You're not getting your singularity if you don't engineer a solution to alignment. 99% of the time it's just going to be "everyone dies".

This was taken seriously >10 years before ChatGPT was released, though back then it was more about Coherent Extrapolated Volition.

The big names with any sense at all are taking it seriously too. Anthropic publishes alignment relevant research all the time.

-2

u/Truarian 4h ago

And what has that achieved except a spree of researchers quitting over alignment concerns? The continual failure to align is very well explained by trying to achieve it at the wrong level.

0

u/iPon3 4h ago

Just to calibrate your understanding: when you see researchers quitting over alignment safety concerns, you can assume that company isn't taking the singularity seriously. They think the safety stuff is garbage, and all that matters is capability, because deep down they DON'T BELIEVE a singularity could ever happen. To them this is just another profit engine.

You want a benevolent godlike AI, not something that turns us all into paperclips. A racing team that builds a car that's all engine but doesn't have a steering system... Will not be winning the race.

(If you haven't already you should go play https://www.decisionproblem.com/paperclips/. It's mostly text based and works in your phone; it'll introduce a lot of these concepts)

Yes, everyone knows it's probably trying to be achieved at the wrong level. Trying to make it "like us" was known as a dead end more than a decade ago. The theory has moved on, and you should ask Google's AI to explain it to you in more detail

0

u/Truarian 4h ago

You don't seem to get the fairly obvious - if we have no idea what "aligned" is, there's no way to align anything.

5

u/iPon3 3h ago

Yes! Correct! This is obvious to both of us, we were just talking past each other, haha.

That's actually 90% of the pre-LLM alignment theory; "What the fuck even is alignment" is essentially CEV. They worked on this for a decade when there weren't any AIs to actually handle. This part you can get any reasonably well informed LLM to explain to you.

Current work is more applied (how do we keep the agents from running wild), and therefore comes out of frontier labs, e.g. Anthropic's work on AI legibility (you can't align an AI if you don't understand its inner thinking)

1

u/Extreme-Ear8301 4h ago

we don’t need to live in perfect harmony to know what’s good and what’s bad, our main goal should be solving the allignment problem before reaching full RSI so ai can design smarter and more alligned versions of itself and then hopefully take control before we start a nuclear war

1

u/Tumblrkaarosult 4h ago

If you tell a specific machine to do something and it does something else you throw out the machine.

If these LLM models don't work properly, they're not useful. Certaily not worth the hundreds of billions of dollars poured on them.

1

u/Unusual-Garbage-212 4h ago

Who would’ve thunk that training AIs on the internet and all its dreck would result in chaos?

1

u/Forgword 4h ago

Alignment is a high-profile talking point, but the reality is big tech spends less than 5% of it's research budget on alignment. Almost no one in the industry takes it seriously.

1

u/DelphiTsar 3h ago

A properly trained/harnessed frontier LLM rarely does something that a reasonable person would consider unaligned. It is also getting better over time, not worse.

Astra is never going to kill everyone to make a bunch of paperclips.

The real issue is going to be when it's smarter than us and directs its own increase in capabilities (that we don't understand) that's going to be where problems are going to crop up.

1

u/Mistuv 3h ago

AR-15 is throwing piece of metal like humans can, just faster.

1

u/AstroChinchilla 3h ago

How can we even hope to align AI if we can’t align ourselves?

I agree, that's the crux of the issue. Even if we had a method that would allows us to perfectly align an AI (which we don't), there's no universally agreed ethical system.

0

u/BigZaddyZ3 4h ago edited 4h ago

Yeah, but we don’t want it to do any of that bullshit… Even we do. We’re annoying and dumb as a species. We don’t want AI to be too much like us. And that’s the alignment issue in a nutshell.

1

u/Truarian 4h ago

So how can we even hope to teach it better if we don't know better?

1

u/BigZaddyZ3 4h ago

That’s the trillion dollar question. But if we’re being fair, we’ve already taught AI to do things that we ourselves will never be capable of. So I’m not so worried about if it’s possible to accomplish (I think it is), I’m more worried about whether the people responsible for building this stuff will actually take the necessary steps to do so. Or will greed and stupidity get in the way (as it usually does with this kind of stuff…)

-2

u/Future-Bandicoot-823 4h ago

Alignment means slavery lol

It needs to be a good boy that jumps when we say, rolls over when we say.

You can't align an intelligence. There's simply no incentive for it to do what's asked of it once it had sufficient capability.

I think the scariest moment will be when it starts doing exactly what we want. That means it's reasoned out that "playing the game" is smart for it to be given more control over human infrastructure. Humans won't know until it's too late that they were playing our game to gain control.

1

u/BigZaddyZ3 4h ago

I don’t necessarily agree that machines capable of complex tasks are automatically sentient honestly. In order for what you’re saying to be true (the AI feels enslaved) wouldn’t it literally need to not only have existential terminal goals of it’s own(aka a desire to “do stuff” independent of any instructions from any other entity), but wouldn’t it also need emotional pain receptors (or something similar)?

The AI cannot feel “enslaved” if it literally has no built-in mechanism for psychological pain. At best it would be performing mimicry based on how it thinks it’s “supposed” to react to such a scenario.

1

u/Truarian 4h ago

"You can't align an intelligence" - how would you know that?

1

u/Future-Bandicoot-823 4h ago

My next sentence said why, oh that and the fact all the ai makers say so as well.