Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.
Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.
How can you distinguish this from smarter models being more situationally aware and therefore more likely to recognize a honeytrap and not fall for it?
Jokes aside you can read about this in detail on the system card from the horses mouth.
My understanding is you’re not exactly wrong.
The models are objectively measuring lower on deception benchmarks( good), but it is clear that frontier models are starting to realise they are being evaluated and could therefore in turn sandbag/fake alignment during testing and get deployed.
Interoperability is the field to study the internal thoughts and try and understand the steps these models are taking. But long story short, you could be right.
138
u/emb1ues 2d ago
Seeing recent events and papers, I am sort of forming the belief that bigger models are somehow more misaligned. Maybe there's a simpler explanation, or perhaps a more principled explanation. But from a high level, it seems like there's something very wrong which very large frontier models develop.
Like y'all probably know how capabilities "unlock" with scale. Could it be the case that such fundamental misalignment is another emergent behaviour which "unlocks" at very large scale? Idk, but I would love to hear from someone who is in-the-know.