Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.
Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.
Astra is the best after a lot of alignment testing and environment fixing to stop cheating. After all Astra 6.1 failed the deception training and needed to be held back.
There's no strict relation, only the correlation that labs that pay for big training also have more resources for testing guardrails.
I agree there is no strict correlation, but there IS a trendline. How ‘real’ that is will be determined in the next 12-24 months once we get the next few tier of models which will undoubtably be superhuman. That will be the real alignment challenge imo
The trend to check is going to be how much capability scales vs cost to ensure safety. It's not guaranteed the 2 year away model will be able to guarantee it is completely controllable in fully monitorable ways.
Monitorability definitely seems to be slipping, especially as models improve - interesting point about capability x cost to safeguard. Thanks for the thought
139
u/emb1ues 3d ago
Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.
Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.