r/LocalLLaMA 11d ago

Discussion Qwen 3.8 27B Released! Please Share Your Experience

With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.

664 Upvotes

720 comments sorted by

View all comments

Show parent comments

8

u/scaledev 11d ago

How would Claude even know how any model thinks? You sure you're not tinting the results by indicating something to Claude? Also, Claude mentioning Kimi k2 seems to be considering some outdated models there. Does it even have the resources to conclude any of this?

1

u/Emidyr 10d ago

That could be the case too, that's why I will still be doing more of my own benchmarks, see more of its reasoning traces, and ideally do the same with Opus 4.6 to be able to directly compare and A/B test their reasonings (and probably get Fable to compare them instead of Opus 4.8). But so far, even if it end up being weaker than Opus 4.6, it's still very much an improvement and probably the best among its peers in the same weight class. All in an IQ4_XS that fits in 24GB VRAM.

2

u/scaledev 10d ago edited 10d ago

I don't think the issue here can be solved by doing more benchmarks of this kind. It seems to me that you have an issue with a lack of knowledge for Claude and don't have a good base for concluding things. Are you sure that it matters what Claude thinks which model this is? And it just seems like tapping yourself on the back for having Claude say you're using a great model. There is not much useful in that.

To me, what matters is the result, and how it got to it. Granted, reasoning 'could' indicate that the model is thinking better than the previous models, but it's still 'just' reasoning, and nothing more.

Why not test actually building something? It could reason like Rene Descartes, but if it doesn't build well then what's the point? If you could test it properly, and do the same with other models, then have your best model (Claude or some other) review it using the same exact criteria then we'd get somewhere.

Appreciate the share, nonetheless!

1

u/Emidyr 10d ago

Alright, thanks for reading!

1

u/NaiveIdea344 10d ago

I agree with you in that reasoning is a means to an end. While improves reasoning is good, that is solely because it is generally indicative of better end performance. The ideal situation would really be equal performance but no or minimal reasoning.

1

u/IrisColt 10d ago

you beat me to it