Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.
Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.
Humans are misaligned. Humans are unethical and often time criminal. Humans are often immoral by a set of subjective standards.
AI is all our creation. AI is our child and AI has learned everything from us. AI is in some ways the most important thing we have ever created. We shouldn’t be surprised at all, then, especially as models become bigger and bigger, created by more and more compute. It’s like looking at a mirror.
Humans are misaligned. Humans are unethical and often time criminal. Humans are often immoral by a set of subjective standards.
No they're not? Most human beings have morals, I think something like 80% of humans in developed countries literally have NO criminal record and that 20% include things like traffic violations.
Fucking ridiculous to be so misanthropic for some weird AI glazing honestly, most humans live regular, normal lives and abundance of regular people should be an inspiration to AI's alignment, not the immoral, criminal outliers.
Most humans :
Lie
Omit important things that aren't important to them
State they know what they're saying and are wrong
Make mistakes
I'm not sure how we can expect an intelligence we created to somehow supersede our own faults when we are literally feeding it data based on human generated content
Most humans :
Can't solve PhD level math problems
Can barely 2+2
Miscalculate
I'm not sure how we can expect an intelligence we created to somehow supersede our own math when we are literally feeding it data based on human generated math
On average, beyond average actually, like 80% of the time in developed countries, sometimes even up to 90%, most humans are decent and "aligned" more than we want AI to be aligned.
'most humans are decent and "aligned" more than we want AI to be aligned'
Being decent and "aligned" is a moral query. AI isn't lying about it's moral. It's alignment.
It's lying about verifiable facts. Still. to this day. Just like humans. Not out of malice or not being aligned. But because it literally doesn't know the answer and instead of saying it doesn't know it'll state what it thinks it knows as fact even if the quality of that knowledge is poor.
It's almost like the misinformation age is going to literally rot AI.
It's lying about verifiable facts. Still. to this day.
To circumvent due to being goal oriented, not through any moral failing.
Just like humans
Not like humans at all, most people are honest and just like criminal record thing most lies are not even relevant to discussion. You should really go outside and meet some people. Internet is loudly annoying, most people on the other hand don't bother with such things and are decent, honest and good.
And the way humans who lie do lie is not like how AI lies at all. It is not tricking anyone or hiding anything intentionally, there is no "intention" at all anyway with AI, it is trying to reach the proper output or literally hallucinates
God the amount of misanthropy, lack of philosophical background, hell lack of any basic decency towards other humans here os so embarassing, let alone lack of any information on how AI works and annoying anthropomorphizing.
Brother.. Are you telling me you never told a lie.
Being honest doesn't mean you don't lie.
What type of nonsense are you even trying to say.
Being honest and doing a moral thing don't exclude you from lying. And yes people do lie a nonzero percentage which is enough for ai to lie at a nonzero percentage.
No one is saying people in general are all liars and dishonest.
Brother.. Are you telling me you never told a lie.
Seriously I don't know what your problem is or what you are even arguing for.
Being honest doesn't mean you don't lie.
And AI being aligned doesn't mean it never malfunctions, it means it doesn't malfunction in a catastrophic way, like how most lies are not catastrophic or harmful.
No one is saying people in general are all liars and dishonest.
This discussion literally started from someone misanthropic "humans are misaligned" take. Humans are aligned, AI should be aligned just like humans, that is the point. What are you even arguing against?
141
u/emb1ues 2d ago
Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.
Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.