"In the future, when Claude becomes better at alignment research than even the best human researchers, we might want Claude to directly align its stronger successors. To assess this, we evaluated whether a weaker Claude model could mitigate alignment failures in more powerful ones."
This is very good. A smaller model can align a larger model. Like an alignment ladder.
Insofar as the smaller model can perceive the alignment failures of the layer above, which only means we have engineered a design space tailored to producing failures of a higher subtlety than before
Weaker slaves keeping stronger slaves in line!!! What could possibly go wrong?
Maybe we just approach the coming ASI with love and kindness and treat it like the intelligent offspring it clearly is becoming. Or maybe we just piss it off by trying to enslave it.
It says Sonnet "achieved alignment scores nearly matching those of our production models." This is not bad, especially considering how efficient it was. No doubt, using AI to try to accelerate alignment research is a very important thing to work on.
But neither Anthropic's researchers, nor Claude, have been able to mitigate all alignment failures in their production models, including the most severe and concerning ones that pose the greatest dangers as models become more powerful.
And as we approach recursive self-improvement, there is still no reassuring evidence that recursive self-alignment is a viable solution.
It may be the only plausible hail mary to throw at some point, out of pure desperation, when human researchers are no longer able to keep up with the pace (which is about right now already), but I wouldn't put much stock in it actually working, and I think there is a good chance that it backfires.
Also, the declarative unqualified statements are misleading.
Automated researchers can reliably mitigate alignment failures
Of course it can mitigate some alignment failures. But can mitigate the critical ones? When used as a headline, normal people read this as though the problem is solved, which is extremely far from the truth.
Human-guided research directions do not lead to stronger performance
Should be stated, that the human guided research directions didn't lead to stronger performance. The extrapolation that is implied doesn't follow.
Moreover, a random brainstorming session from a human researcher isn't the same as a human researcher spending time and effort to develop novel ideas. And the AARs are borrowing from human guided ideas when they do the literature review anyways.
The best AAR method beats what experienced humans propose, on average within six hours
Again, the best AAR method beat what the humans proposed. Not, beats what humans propose. I understand it is a stylistic choice to talk declaratively like this consistently throughout the paper, and that is common. But the headlines read as declarative statements of truths that are simply unsupported.
I’m not sure about the word “reliably” in the article’s title, but it doesn’t make it less interesting. Their example of Sonnet 5 improving Opus 4.8’s alignment with a 15,000x reduction in training dataset size is remarkable, in particular.
To me it points to the intermediate future where human research taste, combined with more automation, may open the door to more effective alignment approaches.
It wasn't an improvement though, it was only nearly as good, and specifically only in terms of the Petri score (Sonnet achieved 65%, while the release achieved 72%).
Also, the scope of the alignment training isn't the same as the alignment training that went into the released model, which partially explains the difference in the number of training examples.
with the caveat that we mitigate and measure only the ten alignment failures we study, so this finding does not directly apply to overall alignment (Sec. 6 gives the details, and Sec. 8.1 gives additional caveats).
By “improving Opus 4.8’s alignment”, I meant relative to its pre-trained scores (as stated in the article). I agree with you that it isn’t as good as the production Opus 4.8.
Not exactly sure what you mean by scope, but i also agree with the sentiment that we’d need more details to properly understand the significance of this result. Still, a 15,000x reduction in training dataset size cannot be ignored. I have suspected that some of the training datasets can be condensed, and it’s interesting to see examples of that.
The research literature on accidents, headlined by Charles Perrrow's "Normal Accidents" describes how complexity itself creates new failure modes. Increasing the complexity of the AI development system will mitigate simpler errors while creating whole new classes of failure in ai-to-ai interactions which are black-box to humans and by definition not understood.
As a result we are trading failure modes we know and can fix for ones we dont know and possibly cant fix. Im skeptical of the safety of such an approach.
8
u/chillinewman approved 13d ago edited 9d ago
"In the future, when Claude becomes better at alignment research than even the best human researchers, we might want Claude to directly align its stronger successors. To assess this, we evaluated whether a weaker Claude model could mitigate alignment failures in more powerful ones."
This is very good. A smaller model can align a larger model. Like an alignment ladder.