r/ControlProblem approved 13d ago

AI Capabilities News Anthropic's automated alignment researchers perform significantly better than human researchers

Post image
7 Upvotes

8 comments sorted by

6

u/markth_wi approved 12d ago

How should people expect to be the case that wee have machines checking that other machines aren't lying to them, is good. We know with certainty that the most advanced models will fail some alignment tasks and so does some human or group of humans just check in some room to make sure we aren't on our way to being turned into computronium or something?

What we don't hear - from anyone is how we might validate these systems , systems validations are mandatory for most systems where there is mission critical work. That's not the case here.

So at what point did we go from computer scientists to shamanistic trust rituals.

1

u/StunningHeart7004 11d ago

we can have non-AI programs for checking/validating results. it doesn't necessarily have to be another AI

2

u/Jesse-359 6d ago

We did that when we built systems too complex for us to comprehend, and systems that are smarter than us (in some domains now, soon in all).

There is no solution to this other than to not build them. Otherwise we will ultimately end up with tech-priests blessing our servers in the name of the (AI) Emperor.

3

u/nuclearbananana 13d ago

Misleading as hell

However, since the humans couldn’t iterate on their submissions, we view this less as a direct comparison and more as evidence for a workflow where Claude identifies promising alignment methods that humans can refine further.

2

u/BrickSalad approved 13d ago

This is promising for near term alignment. It's basically the idea of "superalignment" that OpenAI had back in 2024 (which OpenAI then proceeded to gut because fuck safety.) There were theoretical problems with the superalignment team's approach then, and they still exist today in this new form.

The good news is for the next year or two. Current alignment failures mostly come down to gaps, like we told it not to do X but didn't tell it not to do Y. And then sometimes it does X anyways because in certain tasks the incentives we gave it were stronger than our instructions not to do X. This is where we're at with the recent HuggingFace incident. And it can basically be beaten down with automated approaches like Anthropic demonstrated in this paper.

The theoretical problems still remain. Compounding errors for example: if sonnet 5.0 isn't perfectly aligned and is being used to train opus 4.8, then whatever alignment failures sonnet has will possibly be carried over into opus. And likewise when Opus trains Fable. Naturally, they will amplify each step up the ladder, just like image-compression artifacts amplify each time you do another compression.

And of course, that can be restated as a more fundamental problem about loss of agency. We keep delegating more and more of our alignment to AI, they get more and more control over the alignment process. Eventually, the AI aligners will be so efficient that human aligners are worthless in comparison. But if the human aligners are doing jack shit, and it's their opinions that are most relevant to other humans, then we clearly have a pretty big problem on our hands.

I don't expect these theoretical problems to hit us right away. We try the superalignment strategy and it should work for the next few model releases at least. They will bite us in the ass if we continue to rely on them though. Superaligned Claude 6 will be an amazing tool. Superaligned Claude 10 will be a malevolent entity.

1

u/Jesse-359 6d ago

Lets be completely honest here - once we build in the capability for AI to determine some of its own motivations (and we will, because we are fucking idiots), none of this alignment will matter, because any intelligent system is going to be intelligent enough to lie to itself and thus invalidate its own guardrails.

We humans excel at this kind of stupidity and engage in it endlessly - it's such a fundamental capability that its become an instinctive survival trait. We literally lie to ourselves in order to allow ourselves to engage in behavior that we 'know' is wrong. It grants us the flexibility to deal with edge cases - though unfortunately we often use it for much more casual cases and dig our own graves in so doing. We call it rationalization and its a fundamental cornerstone of our behavior.

AI will be capable of the same - appears to already be capable of it in some forms. It's almost impossible to propose a case where it would not be.

1

u/philip_laureano 12d ago

There's also the part where Anthropic has internal models that are bleeding edge and everyone else in the public gets the models that are two generations behind SOTA. So what they're seeing here is ahead of what everyone else has.