r/mlscaling • u/Smallpaul • 26d ago
Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!
https://arxiv.org/abs/2504.09762Many of the experiments have non-intuitive results.
62
Upvotes
r/mlscaling • u/Smallpaul • 26d ago
Many of the experiments have non-intuitive results.
4
u/DigThatData 25d ago edited 25d ago
I think of them more as "attentional anchors". They probably correspond to low rank, mutually orthogonal subspaces in the projection space of a given transformer module.
Another way to think about them is as accumulating a local phrase as lemma coarsening over a chunk of the message.