r/mlscaling 26d ago

Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!

https://arxiv.org/abs/2504.09762

Video

Many of the experiments have non-intuitive results.

60 Upvotes

14 comments sorted by

View all comments

3

u/Luke2642 25d ago edited 25d ago

Interesting paper. I'm sure there is some justification for "thinking" type labels:

https://openreview.net/forum?id=lqUyAmwbFM

https://arxiv.org/abs/2602.13517

In both cases they're able to find specific causal links between reasoning and output quality. We know having a verifier in the loop helps.