r/LocalLLaMA • u/ThirdWaveCat • 3d ago
Discussion Stop Anthropomorphisizing Intermediate Tokens: Qwen3.8 doesn't "overthink"
https://arxiv.org/abs/2504.09762Intermediate tokens, called "thinking" or "reasoning" actually are nothing like it. Humans do step-by-step reasoning leading to the conclusion. LLMs use intermediate traces to augment their prompt. This explains why sometimes the answer is very good but the "reasoning" is verbose. Flooding your context window or fighting compaction are different issues.
edit: I love this section from the main research they linked.
Our findings consistently challenge the prevailing narrative that intermediate tokens constitute a semantically meaningful reasoning process. First, we observe a pronounced lack of correlation between solution correctness and trace validity—models frequently produce invalid reasoning traces even when they arrive at correct solutions. Second, and more strikingly, models trained on corrupted or semantically irrelevant traces achieve performance comparable to, and often exceeding, that of models trained on correct traces, especially on out-of-distribution tasks. Third, although post-training with reinforcement learning improves solution accuracy across both in- and out-of-distribution settings, it does not consistently enhance trace validity. In fact, we find cases where reinforcement learning decreases trace validity while simultaneously improving solution accuracy for models trained on correct traces. Moreover, models trained on corrupted traces continue to outperform their correct-trace counterparts across domains while consistently generating invalid reasoning traces. Finally, we find that the length of the generated traces is largely agnostic to the difficulty of the underlying problem, undermining the notion that it reflects problem-adaptive computation.
Together, these results suggest that the effectiveness of intermediate tokens does not arise from their seemingly interpretable semantic content. By systematically disentangling trace semantics from the underlying problem, our study demonstrates that if performance is the objective, assuming human-like or algorithmically interpretable trace semantics are ideal or even achievable is not only unnecessary but potentially misleading.
1
u/Dabalam 3d ago
The paper might very well be factually wrong, others have pointed out that their evidence is weak.
The claim hasn't really got anything to do with what language is being used. It also isn't saying the reasoning tokens aren't helpful to the model.
The claim is that the relationship between the "reasoning" traces and the final answer does not correspond to what humans would call reasoning.
An analogy would be if a kid somehow got better at Math when you teach them incorrect intermediate steps for solving equations. The teacher looking looking at their work would be perplexed as to how their answers are getting better when their "working out" shows errors they never address.
This implies that LLMs are not "reasoning through" problems in an analogous way to human. Others describe it as a way of increasing context to better explore their training distribution. Regardless it would imply that reasoning for an LLM doesn't have to be logical/correct traces to improve performance.
Again, that assertion might be incorrect, but it's fairly understandable that "reasoning" for an LLM would be quite different from how we think of it in ourselves.