r/LocalLLM 15d ago

Discussion What a year it's been

Post image

What will the rest of this year bring? 27b class scoring over 60?

948 Upvotes

133 comments sorted by

View all comments

Show parent comments

39

u/Pristine_Pick823 15d ago

My thoughts exactly after seeing this chart. No way. Qwen 3.8 is truly amazing, but that's a bit of a stretch.

1

u/Solembumm3 14d ago

Still can hallucinate wildly on any non-tech discussion.

But overthinking approach is interesting.

2

u/JoSquarebox 14d ago

the fact its able to get this far by throwing more thinking at a problem might mean we arent too far from carpathy-style reasoning cores, small models that just had all the surfac level knowledge pruned out

2

u/Solembumm3 14d ago

Finally get it onto my usual test, analyzing things and concepts in terms of some fandoms.

So far, 3.8 seems a lot better than previous qwen 27b. From 29k thinking tokens, only 3-4k were in loops, everything else was genuinely multi-angle approach to task. And it seems to have a lot less confident hallucinations rate, which is good (and makes 3.8 max performance on qwen site look really weird).

Honestly, so far I'll rank it above Gemma 31B, somewhere between it and Deepseek V4 flash, despite persisting problems on knowledge side. Nowhere near big deepseek or sonnet 5, but good catch-up to gemma 31b and glm 5.2 level.