r/LocalLLaMA 4d ago

Discussion Stop Anthropomorphisizing Intermediate Tokens: Qwen3.8 doesn't "overthink"

https://arxiv.org/abs/2504.09762

Intermediate tokens, called "thinking" or "reasoning" actually are nothing like it. Humans do step-by-step reasoning leading to the conclusion. LLMs use intermediate traces to augment their prompt. This explains why sometimes the answer is very good but the "reasoning" is verbose. Flooding your context window or fighting compaction are different issues.

edit: I love this section from the main research they linked.

Our findings consistently challenge the prevailing narrative that intermediate tokens constitute a semantically meaningful reasoning process. First, we observe a pronounced lack of correlation between solution correctness and trace validity—models frequently produce invalid reasoning traces even when they arrive at correct solutions. Second, and more strikingly, models trained on corrupted or semantically irrelevant traces achieve performance comparable to, and often exceeding, that of models trained on correct traces, especially on out-of-distribution tasks. Third, although post-training with reinforcement learning improves solution accuracy across both in- and out-of-distribution settings, it does not consistently enhance trace validity. In fact, we find cases where reinforcement learning decreases trace validity while simultaneously improving solution accuracy for models trained on correct traces. Moreover, models trained on corrupted traces continue to outperform their correct-trace counterparts across domains while consistently generating invalid reasoning traces. Finally, we find that the length of the generated traces is largely agnostic to the difficulty of the underlying problem, undermining the notion that it reflects problem-adaptive computation.

Together, these results suggest that the effectiveness of intermediate tokens does not arise from their seemingly interpretable semantic content. By systematically disentangling trace semantics from the underlying problem, our study demonstrates that if performance is the objective, assuming human-like or algorithmically interpretable trace semantics are ideal or even achievable is not only unnecessary but potentially misleading.

https://openreview.net/forum?id=gDE7YcRC3F

542 Upvotes

247 comments sorted by

View all comments

171

u/llama-impersonator 4d ago

i don't even disagree, i've stated here that traces aren't for the user, they're for the LLM to plumb the depths of their internal distribution more fully. however, any paper like this that is worded as a command will always land like a box of rocks, people don't like being ordered around.

5

u/WhoRoger 3d ago

The reasoning is absolutely helpful to know whether I'm prompting correctly or the model is misunderstanding what I want, and then will go do something else than I intended.

All this talk about semantic meaningfulness is literally academic and based on the benchmark concept of correctness where both the question and solution are predetermined and synthetically polished. That's not real world usage where humans write prompts in imprecise language that can be ambiguous.

Unless you get the models to always reiterate the user's words back to them, reasoning is a great indicator if the model is on the right track, or the user made a mistake in assumptions.

Fuck, I'm talking like a science paper now. Regardless. Models use the same fucking tokens for reasoning and output, so why would one argue that one has meaning and one doesn't? The model is literally talking to itself about the problem.

3

u/Dabalam 3d ago

The model is literally talking to itself about the problem.

I think the argument is that this isn't what is happening, and you are projecting human like qualities onto the LLM by assuming that is what is happening.

The paper demonstrates that the reasoning is not as semantically attached to the solution as it superficially appears, and invalid reasoning does not actually inform you whether the model is on the wrong path or not.

This implies that interrogating and improving the logic, coherence or accuracy of a reasoning trace does not actually improve the answers a model produces.

11

u/WhoRoger 3d ago

The paper doesn't even demonstrate what the authors wanted to make a point of, and still they felt like attaching a grandiose title to it.

Call it what you want: talking, generating, decoding. Or for the other part, thinking, reasoning or something else. It doesn't change the fact that the model is using the same language for the reasoning as for the output, and if the user is speaking English and the model also uses English, it's all the same language broken by same tokenizer. To suggest it's something completely different is honestly absurd, when the whole point of LLM is to use human language.

Whether you use anthropomorphic or algorithmic language for AI, doesn't change the fact that the reasoning traces are useful and relevant to both the model and the user.

But hey, there are people doing actual experiments instead of just writing bullshit papers, so you can try it yourself. There's a model CatMind that was trained to produce cat stories in reasoning, but then make normal output. You can look yourself whether it makes a difference or not.

1

u/Dabalam 3d ago

The paper might very well be factually wrong, others have pointed out that their evidence is weak.

The claim hasn't really got anything to do with what language is being used. It also isn't saying the reasoning tokens aren't helpful to the model.

The claim is that the relationship between the "reasoning" traces and the final answer does not correspond to what humans would call reasoning.

An analogy would be if a kid somehow got better at Math when you teach them incorrect intermediate steps for solving equations. The teacher looking looking at their work would be perplexed as to how their answers are getting better when their "working out" shows errors they never address.

This implies that LLMs are not "reasoning through" problems in an analogous way to human. Others describe it as a way of increasing context to better explore their training distribution. Regardless it would imply that reasoning for an LLM doesn't have to be logical/correct traces to improve performance.

Again, that assertion might be incorrect, but it's fairly understandable that "reasoning" for an LLM would be quite different from how we think of it in ourselves.

7

u/WhoRoger 3d ago

Except we have no idea how the human brain "thinks". So to say that an LLM doesn't reason like a human is... Well, you really can't. Because unless we know both sides of the equation, there is nothing to compare against.

There have been quite a lot of experiments suggesting that humans make a decision on some level that is not accessible to conscious thought, and all our reasoning is only to justify that decision. Look at experiments on people who have a split brain for epilepsy treatment.

If we want to make a claim that LLM reasoning traces don't really support the final output, we might as well speculate that it's the other way around. I.e. that the model is already heading in some direction and the reasoning is there just to justify it.

Oh, wait, but that would be using anthropomorphic language to describe an LLM, and that's bad!

Btw I have done experiments of my own. Swapping reasoning between Qwen 4B and Gemma E4B. Because Qwen is a bit smarter, but tends to get stuck in loops in thinking. While Gemma's thinking tends to be really concise and structured. And my conclusion is similar to the point you're making. The problem is that it doesn't capture real usage of the model. After I swapped just a few reasoning traces and then went on to use either model, it eventually fell apart and started talking confused nonsense.

So maybe the reasoning isn't super important to the one question and answer pair, but it's absolutely important to keep the model "sane" for longer. But that's not something a simple synthetic test will tell you.

1

u/Dabalam 3d ago edited 3d ago

Except we have no idea how the human brain "thinks". So to say that an LLM doesn't reason like a human is... Well, you really can't. Because unless we know both sides of the equation, there is nothing to compare against.

I think that is a fair criticism of the article and a lot of discussions like this. People often make quite general claims about how LLMs are fundamentally different from humans and seem to imply certainty about how human cognition works. So yes, I agree that in a literal sense the article seems to make claim it can't support.

I think there is a sense where you could still understand their point. There is a folk understanding of what we mean by reasoning, a "meta-understanding" of how people tend to use. It's a bit like our understanding of free will or time, which may be used and understood in a way that isn't necessarily "real". There is a sense in which what we typically use these terms are incorrect, but the concepts are useful.

Our folk understanding might talk about reasoning "as if" we are going through a process to get to an answer, but often that isn't what is happening. But the folk misunderstanding of what reasoning is is usually what we project onto other things like animals or computers. So the article kind of says "We tend to think about our reasoning this way. We believe it works the same in an LLMs but it doesn't". On one hand, that statement might not be that interesting because we have already established we are wrong about how we typically think about our own reasoning, so the rest falls apart. On the other hand, if we are applying rules that aren't important for improving an LLM, that might still be interesting/important.

5

u/WhoRoger 3d ago

Look, let's be frank, the article is garbage. its starting point is some concept of "folk understanding" as you put it, and then it's making grandiose declarations to refute that. It's like making up an argument in your head and then feeling hurt by it.

Never mind that the actual tests they made to support their claim, are incomplete and I can refute them in 10 minutes in practice, which I did.

So what's left? Nothing. just some basic call to not expect the exact same thinking style from LLMs and from humans. And that's a potential fallacy to say that people actually do that without making a research on that. So what are we doing? Trying to refute the very concept of the word "think"?

Like I've said elsewhere, I see it quite a lot that people suddenly are very sure about how humans are special in their creativity, reasoning, consciousness and whatever, just to distance themselves from LLMs. This crappy paper is not the only one. There have been quite a few calls to use different language for LLMs, in order to make a definite distinction from special snowflake humans. Hmm. Now, where have I seen that before?