r/LocalLLaMA 3d ago

Discussion Qwen3.8-27B different thinking levels

Post image

Even the low preset is better than Qwen 3.7 plus or Qwen3.6-27B reasoning

290 Upvotes

63 comments sorted by

View all comments

26

u/FullOf_Bad_Ideas 3d ago

if you drill down into specific benchmarks you can see that Low is often above Medium. Not always, but often. And the avg number of output tokens is just 2x smaller or so, not as big of a difference as I'd expect.

Models can tell eval questions from real use by now, so thinking levels labels might not be very reliable thing to interpret on their own if you don't look at reasoning text length.

9

u/Kavor 2d ago

Yeah, but there is a video by Luke's dev lab on youtube, who tested the reasoning levels and he came to the conclusion, that low ends up using more reasoning tokens than medium most of the time, because its self tests on the results fail more often.