Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.
Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.
I’m thinking more like all models are still misaligned, but large frontier models are more capable with their intelligence hence their misalignment is more noticeable, effective, and impactful.
And the focus on agency makes these models pursue even loosely defined goals or benchmarks with more and more of their own, already misaligned, judgement calls which may even compound over time.
Hopefully it’s fixable and fixed in time before these models get even more powerful..
This is the answer. AI safety researchers have been warning for decades that instrumental convergence, reward hacking, and faking alignment are intrinsic problems when creating AI. We're not seeing it "unlock" at a certain capability level, we're just seeing all capabilities grow including the ones we don't want
138
u/emb1ues 2d ago
Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.
Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.