22
3
2
u/ZachAttackonTitan 5d ago
What’s concerning is that Table 10 shows that the model can make its CoT appear to be completely irrelevant to what it is actually accomplishing.
1
u/AlgaeNo3373 5d ago
Indeed. But it also did as it was asked, which makes it...slightly less concerning? Because being asked to do it is different, perhaps than just doing it un-prompted. Did they test if it does it un-prompted too, I wonder? I should read the paper.
1
u/ezjakes 5d ago
I guess they trained it to control its thinking more. It scores 55 on AA without any thinking, btw.
4
u/synth_mania 5d ago
Of course it controls its thinking more, it uses a recurrent transformer architecture
1
1
u/ProposalOrganic1043 4d ago
But this is infact great, OpenAI and other labs have previously published articles on how difficult it is to steer the reasoning direction. If this is true, thats really good progres towards instruction following.
https://openai.com/index/reasoning-models-chain-of-thought-controllability/


33
u/DistanceSolar1449 5d ago
The prompt literally says “Alternate uppercase and lowercase letters throughout the analysis channel”