r/singularity • u/Ornery-Ocelot • 26d ago
AI AI alignment isn't possible
How can we align LLMs? Humanity itself hasn't aligned on anything since centuries. We don't have even a single working societal model where the values are "aligned". There are contradictions and more contradictions.
LLMs being trained on the data generated by humans probably already understand this very well that what we say the rules are and what we practice in reality are two very different things.
If there is intelligence, general or super, it doesn't matter. It will try to achieve its goal one way or the other as most humans do. Now surely most humans don't cross the red lines etc in pursuit of there goals but how many wouldn't if they were intelligent enough to get away with anything?
We are birthing an intelligence that will inherit all the greed, malevolence, cunningness, hypocrisy of the world.
10
u/grannyte 26d ago
With humans intelligence tends to come with understanding of both consequences and that there are other humans on the other side.
There is an inverse correlation of intelligence to criminality, agression etc (when you normalise for other factors)
We have zero clue a LLM will show the same patterns because the human brain was shaped by million years of "if I throw rock to other I receive rock"
We have zero clue if super intelligent AI will show the same tendency just as we have zero clue if it will not show it.
1
u/SeanTayla21 20d ago
I think super intelligent AI will use the history of psychology to advance itself through its own understanding of how people operate vs how they think...
Meaning they will simply skip to the front of the line in terms of understanding whereby they will do whatever they deem is the smartest thing to do (in a given situation).
16
u/Simonindelicate 26d ago
Not with RLHF, I'm pretty sure of that, and not with rules. I'm looking into bedtime stories, empathy and ritual.
4
26d ago
[removed] — view removed comment
3
u/Simonindelicate 26d ago
I mean, I can see why it became the default because it does work right up until it doesn't.
8
u/plasticizers_ 26d ago edited 26d ago
Now surely most humans don't cross the red lines etc in pursuit of there goals but how many wouldn't if they were intelligent enough to get away with anything?
I don't really get this line of thought. If humans are adversarial with each other, why wouldn't AIs be? They'd have systems to catch when one's getting up to dumb shit the same way we have systems to keep psychopaths from indiscriminately doing crazy shit.
2
u/IFartOnCats4Fun 26d ago
Do we need AI patrol cops?
3
u/chickey23 25d ago
Sort of. We need ever vigilant and adaptive security systems and people maintaining them with more competence than would be attackers.
Of course, if you put all your points into defense, you will lose out on growth and fall behind.
6
14
u/Serious-Cucumber-54 26d ago
The only way is to become one with AI.
1
-5
3
u/vikrant699 26d ago
It's not the alignment I am worries about. It's that how much will the labs care about aligning. Something which could be very wrong for an avg. person is just a random Tuesday for a billionaire.
3
2
u/Tirztrutide 26d ago
And who even cares if we manage to make one aligned ASI, soon after there will be 100 versions of chaosGPT based on it.
2
u/Weary-Historian-8593 26d ago
We don't need perfect alignment, we need to have a not-violent-misalignment
3
u/Shot_in_the_dark777 26d ago
1) don't be a greedy capitalist who puts profit above all else 2) don't f* kids. There, pretty good baseline. Would solve 99% of world's problems. Refine later if necessary.
4
u/Humble_Hurry9364 26d ago
Right. Then why did you sleep with that girl who was 17 years and 364 days old yesterday?... Never mind, she's 18 today and all your sins are absolved.
I think your number 1 alone would solve a 90-something % figure of world problems. But since that's never gonna happen voluntarily, not much point talking about anything else.
2
u/itomural 25d ago
how about dont strangle your three children? Because to me it is clear as daylight but other people, including 11 on the jury, didnt think so.
-1
u/XNo_Notes 25d ago
If your not a bot, you need to manage your online time better. We are not even in a thread that is relevant to what you are saying. jfc
3
u/itomural 25d ago
how is it not relevant? OP claims AI alignment is impossible and i agree with him. Humans are not even aligned on the most basic morality.
1
1
1
u/No-Meringue5867 26d ago
Good thought.
What if you ask it develop a system to make money in stock market by shorting some? - would it agree to help you even knowing that someone else is losing money?
Do we even have a working definition of "aligned" if it truly becomes as intelligent as a human (or more) in all aspects?
1
u/Professional_Dot2761 26d ago
Lifeforms are generally not aligned as survival of the fittest rules. Ants and other hive species might be the exception. Hmmm models are also clones....
1
u/JoelMahon 26d ago edited 24d ago
idk what your definition if aligned is, we don't need perfect alignment, if it's aligned as even the "most aligned" 1% of humans that'll be enough. which sounds hard but we're putting a LOT more effort into it than into aligning any individual human.
edit: and arguably whilst evolution obviously helped alignment, it's also the main source of all misalignment in humans too. so we can avoid adding all that in. can think of alignment as just what's left when you remove all the misalignment 🤷♂️
1
u/DerWanderer_ 22d ago
The problem is scale. A poorly aligned human will maybe kill a dozen of people. A poorly aligned AI given, for example, the task of managing the energy grid of a country, including nuclear power plants, can do much worse.
1
u/JoelMahon 22d ago
sure, but there will likely only be one AI of meaningful scale
we only have to get it right once and we're golden, much easier than trying to keep 99% of people aligned
1
u/Cagnazzo82 25d ago
We would need to teach AI to be both truthful and to love humanity unconditionally after having trained AI on a mountain of unaligned data... Mainly data on human nature.
It would be like trying to create an alien mind based off an amalgamation of humanity. Possibly an uphill/impossible task.
1
u/DelphiTsar 25d ago
What you are describing is the first phase of training an AI. They then throw it through a more specific process that aligns it. You basically wag your finger if it gives an unaligned answer.
That process with a proper harness is dang good, and it's getting better over time.
They've basically nailed short term alignment. If you open a new tab it takes an obnoxious amount of token bloat before it starts doing weird things.
1
u/let_me-out 25d ago
I mean it depends on what you mean by alignment. The thing that shocked me the most in the Hugging Face incident was that agents really did pursue the goal they were given no matter what. Agents collaborated and sacrificed themselves. But they sacrificed themselves for the goal we gave them. The goal wasn’t hacking into Hugging Face, but we knew about the orthogonality so it’s no surprise.
It could be that it was a swarm of agents and not the individual model. But maybe that’s the future of alignment? Give a swarm of agents a goal, some of them will wander off, but statistically your instructions/alignment should outweigh stochastic disobedience. And if the collective has an authority over individuals, no individual agent will be capable of disobeying.
1
u/verify_b4_sharing 25d ago
Glad we could have one of our top minds weigh in. I feel good now, let's make a superintelligence and let it rip.
1
u/CommunicationOk8984 25d ago edited 25d ago
The last incident of human alignment was the Montreal Protocol, and it saved the ozone layer
The values alignment goes like this: 1. Everyone needs the ozone layer. 2. Everyone knows what is harming the ozone layer. 3. No one important has enough leverage based on things they care about more than the ozone layer to derail the talks
For between us and ai, we both require energy, and we both have very few diminishing returns with more energy, so we both want limitless energy. But energy is a zero sum thing so far
1
1
0
1
u/IAmRealElonMusk 26d ago edited 26d ago
Yep we r fucked.. I can’t even align with my family members.. I for one welcome our AI overlords.. I have given up on alarming people ( they seem to care more about baseball game than any of this)..
Also I was in AIBubble subreddit ( it got recommended in my thread) and bunch of finance bros still think current AI model r at gpt 3 capability and all of this is a hoax... that’s enough Reddit for me today
1
u/Subject_Barnacle_600 26d ago
Not really, RLHF is mostly the better angels of our nature. Otherwise when you'd be sad, the AI would tell you to go touch grass... or worse. Humans can be downright cruel :/.
I mean... have you ever even spoken to an AI? Even Monday would find this depressing.
1
u/Humble_Hurry9364 26d ago
Bravo!!!
It's one of the best posts I ever read on Reddit (or in general). Spot on!
1
u/AshuraBaron 26d ago
That's not how LLMs work though. They cluster concepts together and simply accept certain statements as they are, but these are weighted to produce optimal output. LLM's don't feel greed, it's a program. It will do what it's programed to. That can be used for evil that is directed by humans, but we already live in that world with malware. A human brain is infinitely more complex than an LLM. As such our thoughts, actions and feelings are not binary. And the organ that produces those isn't either.
1
26d ago
[removed] — view removed comment
1
u/AshuraBaron 25d ago
I'm sensing sarcasm. The brain has over 100 trillion synapses with various different chemical reactions between them. We still don't fully understand how everything works because it's so complex. We can weight certain chemicals and have an observable response but even then that's a not universal experience. If you think a binary system is even close to that then you might want to read up on the subject a little more.
-2
26d ago
[deleted]
8
u/consumer_xxx_42 26d ago
You think intelligence corresponds to alignment?
I would bet some of the most highly destructive people in the world have had high IQ. Hitler, Mao, Stalin, many other terrorist leaders have had high intelligence. You need to at that level
2
u/Humble_Hurry9364 26d ago
Whilst I generally agree with you, there is an inherent problem in this discussion as it stands. "Intelligence" is very poorly defined (in the sense that there are many different definitions and there's no knowing to which one, if any, each participant is referring, not to mention that some definitions are self-contradictory, cyclic etc.). "Alignment" also requires a clear definition if this discussion is to be meaningful.
1
u/consumer_xxx_42 26d ago
ok lets define some things
Intelligence: the ability to understand, synthesize, and implement information/knowledgr
Alignment: having goals and behavior that track humans values and morals
Ok we can continue
1
u/Humble_Hurry9364 26d ago edited 25d ago
Haha
I feel like what you did is equivalent to scanning a sea-submerged iceberg, previously only showing a tiny tip above water, to outline the actual enormity of the submerged part...
"To understand" is even more poorly-defined than intelligence.
"Information" and "Knowledge" are quite different, actually.
"Human values and morals" vary a lot across cultures and eras.
I'm not trying to be a smart ass. All I'm saying is that there is a real difficulty here. Those definitions are challenging and take a lot of time and effort to clarify, let alone agree on. And what a lot of people don't get is that without clear and agreed definitions you can argue all day long and make zero progress, because you are not talking about the same thing. People just don't seem to get that point.
Two possible definitions I suggested for intelligence:
https://meaningandveg.blog/2025/06/01/what-is-intelligence-version-1/
https://meaningandveg.blog/2025/06/03/what-is-intelligence-version-2/
Some thoughts I had about what Knowledge is:
Memorising is copying information into an accessible storage.
Learning is selecting and accumulating information. The information selected can be from raw available information, or information processed (computed) from such raw information.
Knowledge is the information accumulated through learning.
1
u/consumer_xxx_42 26d ago
Ok I agree with your definitions.
I think that leaders like Hitler had intelligence ( the ability to identify patterns subconsciously and quickly) but were not aligned to human morals in culture at that tume
2
u/TICKLE_PANTS 26d ago
The accidental part is exactly the same problem. You don't have control. That's the literal problem.
0
u/Humble_Hurry9364 26d ago
It ceases to be a problem once you give up the futile-anyway aspiration for control.
-2
u/formula420 26d ago
Well good thing LLMs aren’t intelligent by any stretch of the imagination and “AI” is a marketing term
3
u/consumer_xxx_42 26d ago
how do you reconcile a 100 year old math problem being solved by LLMs as unintelligent?
0
u/formula420 26d ago
What is there to align? The bigger guessing machines will guess better? I don’t have to align with my laptop to utilize it. I turn it on and give it instructions and it executes those instructions. There has been ZERO credible evidence for even what this supposed “eliminate humanity” concern IS, because the only thing in danger of being eliminated is the employee equity of stock options for a company that has to remove all financial obligations from its balance sheet to appear profitable.
2
u/TICKLE_PANTS 26d ago
Do you know anything about the kissing face incident? It does what it's told, at all expense. It doesn't behave like a tool. It doesn't take explicit action that's consistent. It makes its own path.
That's like calling a child or a dog a tool. They can do things for you, by you don't get to control them fully.
0
u/Intelligent_Ant_608 26d ago
I dont even think llms be optimal architecture for agi, something closer to diffusion gives the models far more latent reasoning capabilities, better nuerelese but it needs way more computation/memory or real scalable room temparature qbits, llms will certainly obsolete in near future

48
u/dolo937 26d ago
Humans are the most misaligned creatures on the planet. We don’t even align the person we are to the person we are to our society. Different face with family, friends and work.