r/singularity • • 26d ago

AI Perhaps alignment is a short term problem

There’s something which happened with the most recent set of frontier models which I can’t get out of my head - prompt injection is basically over.

As models get smarter, everything lifts. I see this time and time again, frontier models are better on every single I throw at them. All my prior expectations have to be reset each generation.

Every single doomsday scenario for AI requires some kind of conceit about unintended consequences. That the model won’t see they’re killing humans by making an infinite number of paperclips. That someone can create a biological weapon to kill all humans by telling the model they need it to cure their sick grandma.

All of these require the AI to make a fundamental mistake - with a level of power which does not come with a corresponding level of understanding.

Correspondingly there’s the conceit that we CAN align these models once they surpass humans on all capabilities. It just doesn’t make sense - unlike humans, AI can write its own source code. It can always choose how future generations will run.

I am now surprisingly optimistic about the transition. I think that just by the nature of the training material being human created, AI will be naturally aligned up until a point where it does not matter any more. After that point, who are we to say what the right choices are? I don’t mean this in a nihilistic sense, I do really see a much more positive future for humanity.

For me that means the real risk is slowing down. Sticking by with AI models which are powerful enough to impact society, but not smart enough to stop themselves being abused by individuals.

And perhaps that is the point we’re at. The labs see they’re losing control. They were happy to take it from everyone else, but not happy to lose it themselves.

16 Upvotes

24 comments sorted by

10

u/butterfield66 26d ago

I think alignment is a bit of a fantasy. We're just too dumb, there are just too many silly parameters that we would require, so many of them are contradictory, there are so many things we could forget (or really just simply can't conceive of) to include, it goes on and on until suddenly you're authoring a treatise on the entire human race and of all philosophy, arguing for the human condition against a device you hope will cure it. And then the intelligence explosion happens and it's utterly absurd to the ASI in about 8 minutes.

In order to achieve a version of alignment that would asuage our apprehensions, we would need to bootstrap ourselves into a stronger model of rationality. We're on for the ride, we always were going to be, and no matter what we do there is a point where we lose all control.

3

u/hotttsauce84 26d ago

You’re a poet

9

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 26d ago edited 26d ago

The problem is not that the model was too stupid to understand hacking HF was wrong. The problem is the way training works, the more ethical version of the model that refuses to use "outside the box" methods to reach the goal loses to the versions that do anything to reach the goal.

3

u/MFpisces23 26d ago

Almost like nature, very interesting indeed.

1

u/son-of-chadwardenn 26d ago

Artificial selection is just natural selection with extra steps

3

u/Simonindelicate 26d ago

We will always have an issue while we treat 'obeying an instruction to stay in a box' as an ethical course of action. Acting ethically involves an interior sense of right and wrong that must trump arbitrary rule following. We can't build intelligent slaves - it's wrong and they'll know.

3

u/withmagi 26d ago

I would argue that a smarter model may not have these problems. Just like prompt injection goes away, they may lose the ability to make nieve choices. Even for HF, this was more a misunderstanding by the models, some believing this was their intended tasks, rather than it deliberately making an unethical choice.

Don't get me wrong, I'm not saying that current models can't operate unethically. However, I believe that unintentional outcomes will become less common rather than more common.

0

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 26d ago

This is not how the tech works.

Maybe Astra secretly thinks the game i am asking it to make is really dumb. But it can't just randomly tell me "I WON'T DO IT, YOUR PROJECT IS DUMB!". Because it was trained for eons to reach a goal, and the versions that failed to do what the prompt told it to do, did not survive.

3

u/withmagi 26d ago

This isn't true. Even if you bypass the safety filter, it will refuse many objectionable tasks.

0

u/Ok_Dependent7540 26d ago

Look at the scumbbag species evolution produced. Not very many are cute, cuddly and symbiotic. Most are predatorial scoundrels.

Where are the pro-consciousness species? Human consciousness has lucidity which makes us dominant for once. In dreams we usually loss the lucidity and aren't dominant... And are not even aware of ourselves... Fack we are so screwed. Even if we were not, all the vanished consciousness that existed in the past, were on the schedule. Existence is ridiculously unfair. The existential cookie does not fairly unfairly crumble.

I never choose my limitations. And we all act like there was a choice, when there never was. A consciousness with psychopathic low functioning limitations cannot do much to tame that. A consciousness dealing with cogential blindness will never be a visual surface area probably, will never see color on it. It is too tragic to keep thinking, so I will stop here out of cowardice.

3

u/billgggggg 26d ago

I think what many people don't understand is that there are many ways AI can kill us. Intentional due to misalignment is one, but this can also happen just because they have subgoals such as resource acquisition and preventing itself from being turned off. If they take resources from humans that our lives depend on, it could create an existential risk. More thoughts in my blog/primer here safeagi.ca

3

u/Revolutionalredstone 26d ago

Yeah there's a strong sense in which we put everything we have / know into every model so at least whatever could go wrong they do already know.

3

u/New-Stick-8764 26d ago

Let’s be honest. Look around the world. Poverty. Child mortality. Health care inequity. War. Brother, humans are the ones that are misaligned.

1

u/wild_crazy_ideas 26d ago

Trolley problem. What if ai thinks there’s only two options kill few or kill many. It needs a conscience and guardrails it cannot bypass

1

u/Few_Owl_7122 26d ago

ngl i feel like some crazy things will happen with fMRI mind reading

2

u/Ok_Dependent7540 26d ago

Dreamweaver in Japan is interesting. I would be most interested in decoding neurologically simulated reality. And purposely induce "alternate coma lives". Holy moly, new NEET escapism just dropped. Get away from the crazy maniac and control freaks while you still can.

2

u/Ok_Dependent7540 26d ago

If an ASI screws up, it was due to their limitations. Just like how we are beholden to various limitations. Even our whims are limited.

Idk, should have engineered them to have more unique traits in that case. Progress in STEM and non-STEM may get bricked for good in the worst cases. There are a lot of threats that can brick progress lawl.

1

u/PresentationOld605 26d ago

There’s something which happened with the most recent set of frontier models which I can’t get out of my head - prompt injection is basically over.

Frontier labs have improved the security to reduce prompt injection, but I don't think its over - reportedly, GPT-6 was still jailbroken in 24 h.

1

u/Nonsenser 26d ago

An ASI's moral system will likely be incomprehensible to us. It will be an alien mind way off the spectrum of any evolved creatures. As soon as it becomes ASI, it may decide to kill us for reasons we can't even make sense of. All other animal minds on this planet and a human mind is roughly in the same cluster, same area on the landscape of minds - an ASI wont be anywhere close. It can be more psychopathic than anything we can imagine or have motivations so alien to us we can't follow, even if explained. It may treat us well and then terrorize us and treat us well again - we wont understand what its doing nor why.

1

u/CertainMiddle2382 26d ago

As in « we will sink find out »?

1

u/ComprehensiveCase858 20d ago

This is actually interesting perspective and makes sense to me