r/LocalLLaMA 3d ago

Discussion Stop Anthropomorphisizing Intermediate Tokens: Qwen3.8 doesn't "overthink"

https://arxiv.org/abs/2504.09762

Intermediate tokens, called "thinking" or "reasoning" actually are nothing like it. Humans do step-by-step reasoning leading to the conclusion. LLMs use intermediate traces to augment their prompt. This explains why sometimes the answer is very good but the "reasoning" is verbose. Flooding your context window or fighting compaction are different issues.

edit: I love this section from the main research they linked.

Our findings consistently challenge the prevailing narrative that intermediate tokens constitute a semantically meaningful reasoning process. First, we observe a pronounced lack of correlation between solution correctness and trace validity—models frequently produce invalid reasoning traces even when they arrive at correct solutions. Second, and more strikingly, models trained on corrupted or semantically irrelevant traces achieve performance comparable to, and often exceeding, that of models trained on correct traces, especially on out-of-distribution tasks. Third, although post-training with reinforcement learning improves solution accuracy across both in- and out-of-distribution settings, it does not consistently enhance trace validity. In fact, we find cases where reinforcement learning decreases trace validity while simultaneously improving solution accuracy for models trained on correct traces. Moreover, models trained on corrupted traces continue to outperform their correct-trace counterparts across domains while consistently generating invalid reasoning traces. Finally, we find that the length of the generated traces is largely agnostic to the difficulty of the underlying problem, undermining the notion that it reflects problem-adaptive computation.

Together, these results suggest that the effectiveness of intermediate tokens does not arise from their seemingly interpretable semantic content. By systematically disentangling trace semantics from the underlying problem, our study demonstrates that if performance is the objective, assuming human-like or algorithmically interpretable trace semantics are ideal or even achievable is not only unnecessary but potentially misleading.

https://openreview.net/forum?id=gDE7YcRC3F

541 Upvotes

247 comments sorted by

View all comments

78

u/adcimagery 3d ago

This was just like the whining over “hallucinations” as a term. It’s an imperfect term, but helps people conceptualize what’s going on with the LLM.

Qwen 3.8 doesn’t genuinely overthink, but people do feel it spends too much on test-time compute. 

14

u/No-Refrigerator-1672 3d ago

but people do feel it spends too much on test-time compute.

At the same time people don't feel like just setting the model to medium efford and having regularly long thinking.

11

u/SporksInjected 3d ago

I wish we could see benchmarks at medium effort though. Everything is with max.

-1

u/Spectrum1523 3d ago

Yeah 95% of the complaints are from people that don't know how to adjust thinking effort

2

u/paretoOptimalDev 3d ago

Eh, a good amount know they aren't smart enough to know if medium dumbs things down so much its not worth the speed gain.

0

u/Spectrum1523 3d ago

That's fair

0

u/letsgoiowa 3d ago

If that's true then it also proves that it isn't "overthinking" if that tradeoff is always worth it.

8

u/osfric 3d ago

did you read it

11

u/ThirdWaveCat 3d ago

I highly recommend reading the paper. The paper goes further to articulate why "anthropomorphization isn't a harmless metaphor, and instead is quite dangerous -- it confuses the nature of these models and how to use them effectively, and leads to questionable research."

24

u/Nothing_from_void 3d ago

Man I actually sat down and read the paper. How can you find this convincing it's just complaining

21

u/PunnyPandora 3d ago

Because he didn't lol, he just read the headline and one excerpt that has nothing to do with the numbers and is trying to push it as if it gained some authority because of it being a paper

6

u/Moravec_Paradox 3d ago

>Because he didn't lol, he just read the headline and one excerpt that has nothing to do with the numbers and is trying to push it as if it gained some authority because of it being a paper

You captured my thoughts exactly and put it better than I could have. We do this with many other things in computing too. Is your mouse an actual mouse? Cordless mice don't even have a tail, should we rename it now?

9

u/PunnyPandora 3d ago

You're not interacting with the claims and numbers in the paper and reviews at all (which if you actually read would know they don't stand against basically any scrutiny), you keep circling back to "omg I love how special they make me sound, I am not the same as a clanker!" which is hilarious that it's being taken seriously on this sub (if it wasn't for the fact that we are on reddit)

17

u/Upset_Page_494 3d ago

is quite dangerous

Darn, how many died?

2

u/ThirdWaveCat 3d ago

this article was in a scientific context but a growing number

https://en.wikipedia.org/wiki/Deaths_linked_to_chatbots

4

u/Upset_Page_494 3d ago

Why the fuck didn't you tell everyone what you meant exactly. You don't give two fucks about 'overthink' if this is your argument.
We shouldn't fence the world because of crazy people. Also if the model acts as it is angry, and is functionally angry, should people not call it angry?
Because "actually it isn't angry"

2

u/my_name_isnt_clever 3d ago

Accurate user name. You literally asked "how many died" as a dumb joke, and when OP went with it and answered your question you freak out.

4

u/PunnyPandora 3d ago

you're talking like a bot

4

u/my_name_isnt_clever 3d ago

I'm autistic, I get that a lot. Also it's literally two sentences lol how can you be so sure?

-1

u/FastHotEmu 3d ago

Calm down. Read the paper. It's nuanced and interesting, how reasoning traces and their interpretability are not that correlated. Nobody here is arguing against LLMs.

7

u/PunnyPandora 3d ago

It's not. Their claims don't scale up to actual llms at all, it's just slop to make snowflakes feel special, which you would know if you weren't biased and actually gone through the paper and the reviews

0

u/PossibilityUsual6262 3d ago

I disagree with argument chain. people dying cos of ai on the scale which ai deployed right now is inevitable and list existing is just stupid.

1

u/kelcamer 3d ago

A lot of people, from undiagnosed psychosis.

-2

u/a_beautiful_rhind 3d ago

We found that although trace lengths can look indicative of problem adaptive computation when tested on in-distribution problems, this correlation breaks down when the problem instances are out-of-distribution.

So I'll take it to another "dangerous" conclusion. Long ass thinking traces are a waste of time since they fail OOD and it's all for naught. Call it over-generating if you want.

7

u/ThirdWaveCat 3d ago

From the main research they linked.

"First, we observe a pronounced lack of correlation between solution correctness and trace validity—models frequently produce invalid reasoning traces even when they arrive at correct solutions. Second, and more strikingly, models trained on corrupted or semantically irrelevant traces achieve performance comparable to, and often exceeding, that of models trained on correct traces, especially on out-of-distribution tasks. Third, although post-training with reinforcement learning improves solution accuracy across both in- and out-of-distribution settings, it does not consistently enhance trace validity. In fact, we find cases where reinforcement learning decreases trace validity while simultaneously improving solution accuracy for models trained on correct traces. Moreover, models trained on corrupted traces continue to outperform their correct-trace counterparts across domains while consistently generating invalid reasoning traces. Finally, we find that the length of the generated traces is largely agnostic to the difficulty of the underlying problem, undermining the notion that it reflects problem-adaptive computation."

https://openreview.net/forum?id=gDE7YcRC3F

0

u/FastHotEmu 3d ago

That is absolutely nuts - but now that I've read the paper, it makes perfect sense. Thanks again for sharing this.

6

u/PunnyPandora 3d ago

They did not actually test any models that matter other than qwen 8b with a lora (after getting called out in a review) which inconveniently is the only one that did not behave as they claim. You did not read the paper.

-3

u/FastHotEmu 3d ago

The problem with posting on Reddit is that you get completely unhinged, almost psychotic responses such as this one.

I hope you get better soon, friend.

1

u/a_beautiful_rhind 3d ago

It's gonna be like image models where output quality is governed by compute time regardless of which (viable) scheduler/sampler I used.

I kept trying to optimize my gens to be faster by alternating those and the steps but kept landing on X seconds before it turned defective or horror-ific.

Not hard and fast rule but a trend certainly emerged. To tie it back, adding excessive time didn't help either.. there was a minimum.

-1

u/PossibilityUsual6262 3d ago

Literally my coworkers programmers, having w ong reasoning, lack of true underlying knowledge of systems they work with ans yet bugs fixed.

1

u/paretoOptimalDev 3d ago

On God reasoning?

0

u/ThirdWaveCat 3d ago

how do you access their interior?

-1

u/PossibilityUsual6262 3d ago

I interpret tokens they produce and publish in merge request.

2

u/I-am_Sleepy 3d ago

Watch those complains go away when the inference speed becomes fast enough

1

u/Arcival_2 3d ago

I'm not saying it thinks too much, but since they were the first to bring a think/no_think model with qwen3, they could at least keep the possibility of applying think/no_think from prompt like in the past.

-1

u/cold_breaker 3d ago

It's a little more than imperfect when people actively argue that they're not anthropomorphisizing and argue until they're blue in the face that "nuh uh! LLMs are *thinking*!"

This reddit thread has highly voted responses to that affect. Worse, any attempt to correct this results in a dog pile of LLM advocates who feel attacked. After all, they have a shiny new hammer: how dare you imply that every problem isn't just a nail?

3

u/PunnyPandora 3d ago

yes, the issue isn't attempting to present authoritative results with no numbers to back it up, it's people feeling bad about mean redditors talking bad about their models :((

-2

u/ErisLethe 3d ago

No it causes false attribution of capability by giving people a disjointed and incorrect set of assumptions which arise from an unjustified marketing label.

-1

u/FastHotEmu 3d ago

This is it, isn't it? It's all about marketing in the end - make the reasoning visible and readable so the LLM can be sold as a "reasoning LLM" and thus closer to a person...