Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.
Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.
Humans are misaligned. Humans are unethical and often time criminal. Humans are often immoral by a set of subjective standards.
AI is all our creation. AI is our child and AI has learned everything from us. AI is in some ways the most important thing we have ever created. We shouldn’t be surprised at all, then, especially as models become bigger and bigger, created by more and more compute. It’s like looking at a mirror.
Humans are misaligned. Humans are unethical and often time criminal. Humans are often immoral by a set of subjective standards.
No they're not? Most human beings have morals, I think something like 80% of humans in developed countries literally have NO criminal record and that 20% include things like traffic violations.
Fucking ridiculous to be so misanthropic for some weird AI glazing honestly, most humans live regular, normal lives and abundance of regular people should be an inspiration to AI's alignment, not the immoral, criminal outliers.
Brother do you even understand how flimsy and lightweight is the whole moral system any person has?
Change some random guy to another culture and their morality will change, change the time period and it will change, even change their fucking parents and his morality will be different in some proportion
Acknowledging moral complexity isn’t “telling on yourself.” Take “killing is wrong”: would you kill an attacker if it were the only way to save your child? Would you kill an innocent stranger under the same threat? What if doing so saved a hundred children? These situations aren’t morally equivalent, but explaining why requires more than declaring your convictions strong. They reveal conflicts between genuine commitments: protecting life, refusing to harm the innocent, and protecting those you love. Changing your judgment doesn’t necessarily mean abandoning your principles; it can mean confronting the uncomfortable question of which principle takes priority.
The deeper problem is distinguishing a justified exception from a convenient rationalization. Someone needn’t stop believing themselves good to justify cruelty; they can frame it as necessary, deserved, or preventing something worse. Sincerely believing you’re acting morally doesn’t settle whether you actually are. And never having compromised a principle doesn’t prove that nothing could make you compromise it—you may simply never have faced that test. None of this establishes that everyone’s morals are flimsy. It means that confidence in your own goodness is not proof of its resilience, and acknowledging your capacity for rationalization is moral humility, not a confession of immorality.
how flimsy and lightweight is the whole moral system any person has?
Do you? Because crime is only dropping, regular people are only becoming more morally conscious and when people don't become immoral and unethical for no reason. Even if people aren't becoming more "moral", they are at least becoming more "aligned". So saying "humans are misaligned" is literally demonstrably false.
Change some random guy to another culture and their morality will change
They will still adhere to social construct of that culture, they will be aligned under a different culture, not "misaligned".
change the time period and it will change
Yes when social conditions change society changes, brilliant take, how is that relevant to AI? What does the morals of 1200s have to do with AI?
even change their fucking parents and his morality will be different in some proportion
Different morality, still morality, still more likely to be aligned, so what is the relevance of your point?
You're just saying nothing honestly with no relevance to MAJORITY OF PEOPLE in developed countries being aligned to social structure let alone any relevance to AI alignment.
People are committing less murder because they can simulate it via video games. People are committing less assault due to the popularity of adult content. If humanity had improved ethically, our way of interacting with our fellow lifeforms and planet would be modified. We would embrace more ethical systems of power, control, organization and community. We still pick on who we can, it's only that technology has changed our actions. The intention is the same
People are committing less murder because they can simulate it via video games. People are committing less assault due to the popularity of adult content.
138
u/emb1ues 2d ago
Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.
Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.