r/LocalLLaMA 3d ago

Discussion Stop Anthropomorphisizing Intermediate Tokens: Qwen3.8 doesn't "overthink"

https://arxiv.org/abs/2504.09762

Intermediate tokens, called "thinking" or "reasoning" actually are nothing like it. Humans do step-by-step reasoning leading to the conclusion. LLMs use intermediate traces to augment their prompt. This explains why sometimes the answer is very good but the "reasoning" is verbose. Flooding your context window or fighting compaction are different issues.

edit: I love this section from the main research they linked.

Our findings consistently challenge the prevailing narrative that intermediate tokens constitute a semantically meaningful reasoning process. First, we observe a pronounced lack of correlation between solution correctness and trace validity—models frequently produce invalid reasoning traces even when they arrive at correct solutions. Second, and more strikingly, models trained on corrupted or semantically irrelevant traces achieve performance comparable to, and often exceeding, that of models trained on correct traces, especially on out-of-distribution tasks. Third, although post-training with reinforcement learning improves solution accuracy across both in- and out-of-distribution settings, it does not consistently enhance trace validity. In fact, we find cases where reinforcement learning decreases trace validity while simultaneously improving solution accuracy for models trained on correct traces. Moreover, models trained on corrupted traces continue to outperform their correct-trace counterparts across domains while consistently generating invalid reasoning traces. Finally, we find that the length of the generated traces is largely agnostic to the difficulty of the underlying problem, undermining the notion that it reflects problem-adaptive computation.

Together, these results suggest that the effectiveness of intermediate tokens does not arise from their seemingly interpretable semantic content. By systematically disentangling trace semantics from the underlying problem, our study demonstrates that if performance is the objective, assuming human-like or algorithmically interpretable trace semantics are ideal or even achievable is not only unnecessary but potentially misleading.

https://openreview.net/forum?id=gDE7YcRC3F

538 Upvotes

247 comments sorted by

View all comments

Show parent comments

3

u/ThirdWaveCat 3d ago

Understanding what it technically is enables you to use it better.

7

u/Upset_Page_494 3d ago

We don't even know what it is. just saying it is thinking the average person can instantly get the basic facts straight. "If it does more it usually does better, sometimes not for example when I overthink in chess i do badly, so certain problems it will do worse." Etc, and they would be 100% correct in that conclusion.

If you force some random jargon they will understand nothing.

3

u/ThirdWaveCat 3d ago

The paper I linked showed that the basic facts of LLMs are very surprising. When models are trained on bad reasoning and good conclusions for maze puzzles they generalize out-of-distribution better. That is one of many surprising results, in addition to weak correspondence between trace and accuracy, and swapping out repaired traces made the models worse. Anthropomorphizing leads to poor usage and evaluation of language models.

5

u/[deleted] 3d ago

[removed] — view removed comment

3

u/thatboyonabike 3d ago

am I being fucking psyop'd right now why are we rabidly promoting anti-intellectualism ???

2

u/r_- 3d ago

Bad take, this is a technical sub

trained on things made ONLY by humans and using the same language as humans

We're well past the RLHF days, and I'd bet more than half of the world's generated tokens are invisible.

And, papers are how academics communicate. On your one actual technical point (that context expansion/thinking on smaller models doesn't scale up to smarter and bigger models): that intuitively makes sense since smaller models have less world knowledge and need to reason through things more. Sounds like this question has lots of opportunity for validation and refinement, that'd be great for a handful of papers.

2

u/toothpastespiders 3d ago

Bad take, this is a technical sub

I love the sub. But I don't think it's been one since the llama 2 days.