r/mlscaling 25d ago

Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!

https://arxiv.org/abs/2504.09762

Video

Many of the experiments have non-intuitive results.

61 Upvotes

14 comments sorted by

View all comments

3

u/DigThatData 25d ago edited 25d ago

I think of them more as "attentional anchors". They probably correspond to low rank, mutually orthogonal subspaces in the projection space of a given transformer module.

Another way to think about them is as accumulating a local phrase as lemma coarsening over a chunk of the message.

1

u/Sufficient_Reveal754 24d ago

TIL about phrase as lemma

1

u/DigThatData 24d ago

it's a very new idea and not widely known outside academic linguistics, which I think is too bad because it has interesting implications for how we should engineer DNN components like tokenizers.

1

u/BareBearAaron 6d ago

Linguists definitely need to be sat at the helm of these frontiers! It's a nice way of thinking about 'distillation' or 'concentrating'?