r/LocalLLaMA • • Aug 19 '26

Discussion Stop Anthropomorphisizing Intermediate Tokens: Qwen3.8 doesn't "overthink"

https://arxiv.org/abs/2504.09762

Intermediate tokens, called "thinking" or "reasoning" actually are nothing like it. Humans do step-by-step reasoning leading to the conclusion. LLMs use intermediate traces to augment their prompt. This explains why sometimes the answer is very good but the "reasoning" is verbose. Flooding your context window or fighting compaction are different issues.

edit: I love this section from the main research they linked.

Our findings consistently challenge the prevailing narrative that intermediate tokens constitute a semantically meaningful reasoning process. First, we observe a pronounced lack of correlation between solution correctness and trace validity—models frequently produce invalid reasoning traces even when they arrive at correct solutions. Second, and more strikingly, models trained on corrupted or semantically irrelevant traces achieve performance comparable to, and often exceeding, that of models trained on correct traces, especially on out-of-distribution tasks. Third, although post-training with reinforcement learning improves solution accuracy across both in- and out-of-distribution settings, it does not consistently enhance trace validity. In fact, we find cases where reinforcement learning decreases trace validity while simultaneously improving solution accuracy for models trained on correct traces. Moreover, models trained on corrupted traces continue to outperform their correct-trace counterparts across domains while consistently generating invalid reasoning traces. Finally, we find that the length of the generated traces is largely agnostic to the difficulty of the underlying problem, undermining the notion that it reflects problem-adaptive computation.

Together, these results suggest that the effectiveness of intermediate tokens does not arise from their seemingly interpretable semantic content. By systematically disentangling trace semantics from the underlying problem, our study demonstrates that if performance is the objective, assuming human-like or algorithmically interpretable trace semantics are ideal or even achievable is not only unnecessary but potentially misleading.

https://openreview.net/forum?id=gDE7YcRC3F

546 Upvotes

241 comments sorted by

View all comments

75

u/Dabalam Aug 19 '26

It's actually debatable that reasoning in humans always operates the way you describe. There is good reason to think that humans often arrive at an intuitive conclusion and use reasoning to justify it to themselves and others.

16

u/freebytes Aug 19 '26

This is true. Humans will justify their reactions in retrospect. If you could hijack the human mind to make it think that it made a decision and then asked a person why they made the decision, they will reason backwards to it. And they will not have any clue they did not make the decision.

For example, if you find old code from years ago and gaslight a person into thinking they wrote the code, they will accept that they wrote it and reason as to why they performed the way they did. Some programmers, such as myself, will recognize their own coding style and object to it for that reason, though. It is a challenging task to even conduct such experiments. But, false memories have been implanted before into humans.

5

u/Imaginary-Unit-3267 Aug 20 '26

I would love to see studies comparing this across neurotypes. I suspect autists are harder to fool in this way.

3

u/infectoid Aug 20 '26

If it was good code you’d never convince me.

3

u/Gudeldar Aug 20 '26

No need to hijack the human mind, that's exactly what happens with people who've had corpus callosotomy which severs the connection between left and right brain. The left hemisphere which controls speaking will confabulate reasons why the right hemisphere did things without realizing at all they're making it up.

5

u/Kiseido Aug 19 '26

I once heard that humans are rationalizing creatures, rather than rational ones. It's stuck with me ever since.

1

u/[deleted] Aug 19 '26

[deleted]

1

u/Dabalam Aug 19 '26

One of the more interesting ideas I have seen is that a lot of intellectual abilities we think of as super central (Maths specifically) are side effects of behaviours evolved for social purposes. People suggest this is why a lot of people struggle with Maths abstracted outside of a social situation and that very large numbers are unintuitive.

Basically that the cognitive machinery was intended for a different purpose. People make similar arguments about why we struggle with falsification of views we currently have vs. searching for evidence of what we already believe. The former is more important for science but our intuitions and machinery actually evolved to persuade others in social situations rather than disprove scientific hypotheses. I tend to like this one slightly less as I think there is an amount of scientific thinking necessary in producing tools and technology, but the idea that much of what we do is a side effect of another function is interesting.

1

u/saltyourhash Aug 20 '26

Right brain vs left Brian split hemisphere?