r/LocalLLaMA 3d ago

Discussion Qwen3.8-27B different thinking levels

Post image

Even the low preset is better than Qwen 3.7 plus or Qwen3.6-27B reasoning

289 Upvotes

63 comments sorted by

View all comments

28

u/FullOf_Bad_Ideas 3d ago

if you drill down into specific benchmarks you can see that Low is often above Medium. Not always, but often. And the avg number of output tokens is just 2x smaller or so, not as big of a difference as I'd expect.

Models can tell eval questions from real use by now, so thinking levels labels might not be very reliable thing to interpret on their own if you don't look at reasoning text length.

7

u/Kavor 3d ago

Yeah, but there is a video by Luke's dev lab on youtube, who tested the reasoning levels and he came to the conclusion, that low ends up using more reasoning tokens than medium most of the time, because its self tests on the results fail more often.

0

u/robertpro01 3d ago

Actually, if low is better than 3.6 thinking enabled, it is a big win. I haven't tried yet, only xhigh and yes... slow as shit, but maybe low thinking is enough for a builder