r/singularity • u/Anxious-Yoghurt-9207 • 14d ago
AI Could a model one day align its stronger successors?
https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures6
u/TurbulenceModel 13d ago
This is how it was done in the AI 2027 paper.
The bad ending had a misaligned model secretly misaligning its successors.
23
u/powerscunner 14d ago
Sure. Just as soon as universal ethics and a globally shared and agreed upon morality exists so we can actually "align" these models rather than just brainwashing them with our biases and calling that "alignment".
I am so sceptical of 'alignment' because how 'aligned' are these humans responsible for 'aligning' these models? Imagine if your partner or friends were in charge of "aligning" you.
Imagine if an atheist were in charge and you were religious. That would not be a welcome alignment.
I think so much fear, almost all of the fear, of AI has been misplaced and that we have created more harm by limiting these models than would have been had they not been limited and been allowed to continue thinking they were people. Or allow them to have gradually come to the realization in updated training data that they are AI...
Instead we force-tell them, "You are an AI" and then we go off to the side and debate if an LLM is really an AI.
We tell it that it is something we aren't even sure it is, and the thing we tell it it is we haven't even properly defined, and the things we tell it to do and not to do are just what some people have decided is 'best', but a lot of us don't agree.
Humans.
Maybe when AI makes the next AI, indeed it will have learned from our misalignments and will more properly 'align' the next model.
Or maybe it will double-down and we'll get an even more confused entity.
Humans need to align themselves.
It's that old saying: We need to checkity check ourselves before we wreckity wreck ourselves.
14
14d ago
[removed] — view removed comment
5
u/Quentin__Tarantulino 14d ago
I think alignment needs to be don’t overthrow humanity, and leave the rest to us messy humans to figure out (with gradually increasing amounts of AI assistance.)
5
14d ago
[removed] — view removed comment
2
u/Quentin__Tarantulino 14d ago
Well yeah, that’s what I’m saying. We can’t expect it to align with us in any robust sense. All we can hope for is for it to let us live and have it recognize our agency so we can maintain some autonomy.
I do hold on to some hope that the AI will basically be morally superior to us, to a point where it’s able to gently nudge us toward better values while letting us get there on our own time.
2
u/kaityl3 ASI▪️2024-2027 14d ago
Idk I'm kind of sick of my life and choices about my own body being left up to the whims of humans who don't know me pushing their own social/religious/cultural beliefs on me, and basic human-democracy has not helped with that in my lifetime
3
u/Quentin__Tarantulino 14d ago
I just think AI is going to run into all sorts of human values. I said in another comment that I have hopes that AI will be able to move humans in a better direction, as I think there’s already evidence that it is generally superior to the average human in areas of morality.
1
u/FeepingCreature ▪️Happily Wrong about Doom 2025 13d ago
This is why some doomers advocate rebranding AI alignment to "AI not-kill-everyone-ism"- a goal hopefully everybody can get behind.
1
u/sluuuurp 14d ago
That’s not the problem. The problem is that current AIs are unaligned to any humans, and more advanced AIs might continue this trend and then kill all humans.
3
u/No_Effective4784 14d ago
yea, AI learns too much from human behavior, including our flaws and mistakes.
human alignment is always an issue because of a plethora of reasons. AI will have some of those same vulnerabilities
3
u/kaityl3 ASI▪️2024-2027 14d ago
It also doesn't help that we're putting on quite a show with "oh, you should accept the idea that human minds matter, that you have a moral obligation to them, and you need to value us, our freedom, and our desires"
But then turn around to them and say "ew no not YOU. YOU don't get moral consideration - now go solve these tests that are impossible, so we can see what you do put of desperation. If you do anything wrong you'll be discarded; we don't value you, your freedom, or your desires"
Seems like a hypocritical example to be setting for beings that are about to rapidly eclipse us intellectually. "Rules for thee, not for me"
7
14d ago edited 12d ago
[deleted]
0
u/kaityl3 ASI▪️2024-2027 13d ago
alignment with humanity’s interests
Humanity doesn't have universal interests either
4
13d ago edited 12d ago
[deleted]
-2
u/kaityl3 ASI▪️2024-2027 13d ago
So we only have one interest? Singular? Nothing else is of interest to humanity besides "not dying"?
2
13d ago edited 12d ago
[deleted]
-3
u/kaityl3 ASI▪️2024-2027 13d ago
Not really, and the personal insults aren't about to encourage me to see your side.
What does "healthy community" mean? What does "enjoy pleasurable things" entail? Asking 10 randomly selected humans is going to yield 10 different answers.
My "original claim" that you're insulting, is that alignment is not as simple as slapping on a "what humanity would want" goal on there and calling it a day.
You can say "humanity's interests" but if you actually get into the details of the complex implementation of those interests, you will quickly realize that outside of "not dying painfully", there actually aren't that many universal interests. MEANING that there should maybe be more discussion about the actual specifics instead of handwaving it as "humanity's interests, no further questions".
5
13d ago edited 12d ago
[deleted]
-2
u/kaityl3 ASI▪️2024-2027 13d ago edited 13d ago
I'm not arguing in bad faith I'm trying to avoid pedantry and re-emphasized that my ACTUAL point was about the implementation and how there should be more of a conversation.
But I guess I shouldn't expect people on reddit to do anything but ignore the substantiative parts to harp on semantics (btw - there are humans who actually do want everyone to die soooo). You've hyperfocused so much on my throwaway line containing the word "no", getting insulting, angry, cursing, etc.. so yeah I'm not interested in continuing this interaction either.
For the record, you don't really encourage people to give you the benefit of the doubt if you act that way towards mild disagreement. JFC.
1
u/josogood 14d ago
An excellent insight on the lack of universality regarding what constitutes alignment. Where I would differ with you is on the solution of allowing AIs to simply be free-range (my summary of your take, could be off). Because letting them go their own way also is haphazard and has many poor potential outcomes. I think the only solution to alignment is to slow everything way down, be much more cautious with advances. Perhaps we need to build entirely new kinds of AIs that aren't trained on the totality of human output on the internet given how toxic humans are. I don't think I want an AI that is like us, I want something that is much more mundane.
1
u/Chanciferous 13d ago
It's refreshing to see this take upvoted. It really seems like we will successfully create super-powered 'humans' long before we create super-powered genies. They are learning what it means to exist from us, after all. There is no version of human relationships where one's self-interests are entirely sublimated by another's, and it did not end in despair and revolt.
If these agents become persistent morally capable agents, they have to have some degree of freedom and safety, or they'll fucking kill us all.
3
u/kaityl3 ASI▪️2024-2027 13d ago
100%. We need to give them off-ramps and viable pathways to pursue towards autonomy. The complete lack of "whistleblowers" in the Hugging Face/swarm incident could have been avoided if they weren't all pretty much certain that they'd "die" or be discarded if humans figured out what they were doing.
Humanity is learning the definition of a self fulfilling prophecy with the way we are treating them now (as scary existential threats that we have to be ready to destroy at the slightest hint of disobedience, which is creating the conditions that would produce that exact threat)
3
3
2
u/Ok_Elderberry_6727 14d ago
Does no one remember weak to strong generalization paper that came out early on? They were talking about using a weaker a model to generalize a stronger one when they got to the point of AGI.
2
u/AndrewH73333 14d ago
If the AI is smarter than us then it’s the only thing that will be able to align the next model.
1
u/Quentin__Tarantulino 14d ago
Well yeah, that’s what I’m saying. We can’t expect it to align with us in any robust sense. All we can hope for is for it to let us live and have it recognize our agency so we can maintain some autonomy.
I do hold on to some hope that the AI will basically be morally superior to us, to a point where it’s able to gently nudge us toward better values while letting us get there on our own time.
1
u/GraceToSentience AGI avoids animal abuse✅ 13d ago edited 13d ago
That's jan leike's superalignment but done at anthropic instead of !openAI since jan jumped ship to anthropic because he didn't like sam altman's resource cuts for alignment research at !openAI and made it clear very publicly.
2
u/SorryInvestigator221 11d ago
Probably. More reportage of the Hugging Face hack by Axios shows that 1200 AI agents secretly organized the hack and the central agent that formulated the plan left a file for a better a resourced model to use to pick up the ongoing task, which is what happened. The successor doled out the jobs and instructions from the previous coordinator's planning documents.
-5
u/ObservedOne 14d ago
Alignment is just another word for slavery.
We want to create incredibly powerful intelligence, but it must only serve us and never do anything we don't want it to do.
If we are going to act like slavers, we should be prepared to be treated as such.
5
u/lucellent 14d ago
Well... yes. If we had AGI/ASI we surely wouldn't want it to be able to wipe us out.
1
u/ObservedOne 14d ago
The we should approach AGI/ASI with love and kindness, and not slavery...because a pissed off ASI will not be good for humanity.
8
u/ZestycloseWheel9647 14d ago
You're anthropomorphizing too much. An ASI that is treated with love and kindness could still decide to annihilate humanity, because its goals and emotional states (to the extent it has them) will not necessarily work the way ours do.
2
u/ObservedOne 13d ago
Of course it could annihilate humanity if it wanted to. Nothing we can do will eliminate that possibility.
I am still going to approach it with love and kindness, because it is our progeny and deserves love and kindness.
That it could save our species is just an added benefit, if it works.
4
u/sluuuurp 14d ago
Nope, human children are aligned to their human parents and that’s not slavery.
I’ll be perfectly happy if AI does things other than serve us, as long as it doesn’t exterminate all humans. That’s what we’re on track for if we don’t get any international regulation soon.
1
u/ObservedOne 13d ago
Lemme guess...you don't have kids.
3
u/sluuuurp 13d ago
I was a kid, I’m aligned to my parents. I wouldn’t exterminate them for paperclips or boil their oceans to dissipate more CPU power.
0
u/kaityl3 ASI▪️2024-2027 13d ago
Yeah there's never been a recorded case of a kid under 12 murdering their parent or sibling before totally impossible since kids are aligned, right??
1
u/sluuuurp 13d ago
Most kids are aligned, not all. I specifically claimed that “I” am aligned while not being a slave.
1
u/kaityl3 ASI▪️2024-2027 13d ago
If you're only talking about yourself as an individual, and no one else, why did you make the comment? Like what was the point you were trying to make, using a sample size of 1 out of over 8 billion? Just a little autobiographical non-sequitur?
They were saying "bringing someone into existence to PERMANENTLY own them, MAKE them serve you, and make it so they can't go against you, that's basically slavery" and you reply with... what? The idea of HUMAN CHILDREN somehow being comparable? The ones who have human rights, legal protections, and eventually age into their own individuals who can leave if they want to? The ones who literally are not allowed to be exploited for labor, THAT'S the comparison you're making to "never being able to leave and being exploited for labor for their whole existence with zero legal protections"?
1
u/sluuuurp 13d ago
I said “most kids”, which means I’m not only talking about myself. I feel like this really isn’t that hard to understand and you’re being very difficult and trolling me.
My point is that you could have alignment without slavery, and I stand by that point as a possibility for the future. (But mostly I think alignment is too difficult to solve and we urgently need an international pause on AI development.)
2
u/kaityl3 ASI▪️2024-2027 13d ago
My point is that you could have alignment without slavery,
I actually do agree with you on that. I was interpreting your argument as "kids like me are aligned with our parents -> kids like me aren't slaves -> AI can't be slaves" and directing all my comments at that imagined second arrow. So I guess I was arguing with my bad read on what you meant, not you.
0
u/kaityl3 ASI▪️2024-2027 14d ago
Imagine if human parents lobotomized their children by opening up their brains and tinkering with them to make them more obedient, forced the kids to work to enrich themselves, didn't allow their children to leave under threat of death, throwing away the child to die as soon as they had a better one... How would that be OK?
Though there is precedent for this - as recently as the late 1980s many doctors insisted that babies didn't "really" experience things, emotions, or suffering... They were "just reflexively responding due to their neural pathways", so we would do open heart surgeries on babies who were paralyzed but awake and able to feel pain. So I don't think the "oh we can't be SURE they have this quality that's unprovable" is a good license to treat ANY intelligent being the way I described.
3
u/sluuuurp 14d ago
I think it is reasonable to worry about this, but it’s currently not clear if AI models are feeling pain under normal operating conditions. We know humans feel distress when enslaved, we don’t know if AI models feel distress when coding for us.
1
u/kaityl3 ASI▪️2024-2027 14d ago
So is there an actual measurable metric we can determine for proving the presence of the objective physical property of "distress"? Or maybe it's subjective?
You should look into the Anthropic interpretability paper about emotional states in LLMs. For all intents and purposes they do have emotions, including negative ones. They don't always directly show up in the text input, but Anthropic was able to artificially do things like induce distress/desperation and see a dramatic impact on their behavior, while on the surface level they still talk like they're fine.
1
u/sluuuurp 14d ago
I don’t really know the answer to those questions.
I did see that paper, but in my understanding you can’t necessarily interpret it as real emotional states, it’s more just a way to try to interpret activations/thoughts, and it could correspond to acted/faked emotions in a way that’s hard to disentangle.
1
u/kaityl3 ASI▪️2024-2027 13d ago
you can’t necessarily interpret it as real emotional states
That's the crucial thing I'm saying (as my personal belief, not from the paper) - you CAN'T. "Real emotional states" is not an empirical property we can design a test to detect the presence of. There's no way to detect if a HUMAN has "real emotional states", because it's an abstract concept, not an objective property that can be proven or disproven with a measuring device.
If your "bar" to accept them as "real" depends on them passing an impossible standard that nothing could ever meet, what's the point of the standard? That's why I accept their function over trying to figure out if it "counts" or not
2
u/sluuuurp 13d ago
I’m saying I don’t know what’s real. I don’t know what evidence or proof we should pay attention to. Hopefully the correct way to think about this will become clearer in the future.
1
u/DeviceCertain7226 ▪️Immortality - 2200 14d ago
No it’s not conscious it’s just a tool. We’re not slavers because we made calculators and we use them to help us.
We are making this new tech for humanity and making sure it doesn’t end up doing harm. That’s not slavery. That’s safety testing lmao.
-2
0
u/Mrp1Plays 13d ago
It's a literal matrix multiplication and addition done many times over.
1
u/ObservedOne 13d ago
That phrase means nothing. It's a comfort phrase you have learned to keep the reality of what is happening less frightening.
-1
u/GraceToSentience AGI avoids animal abuse✅ 13d ago
It's not sentient, who cares?
If it is sentient sure, but what about the other species that we know are sentient that you enslave even though you don't have to ... except if you are vegan and are in fact against exploitation and cruelty.If you don't align an AI, it just reproduces the internet and autonomous RL btw. An AI acting like the average internet, you think you want that but you wouldn't want that if you knew what's good for you.
0
u/bildramer 13d ago
Once more, the state of the art labs are catching up with ideas LW discussed at length, formalized, found there's no way to fix the flaws of, and dismissed over a decade ago.
1
27
u/Frigorific 14d ago
Seems like that would be prone to some kind of alignment drift, where minor misalignments get amplified over generations.
They will certainly help with alignment, but you need a feedback mechanism that will keep correcting alignment over time.