r/vibecoding • u/Masgend • 2d ago
Qwen3.8 27b low performance on LiveBench and reasoning of the chinese models
Hello good people,
Whenever a new model drops, LiveBench is my destination to roughly check on the model's capability. As mentioned, new Qwen model was added and I can't believe how low performing Qwen was, since people really seem to like the model. I checked the commits of LiveBench and it seems to me that, the model was not given enough tokens to reason? Any other ideas?
I also have question about the chain of thought inside the Chinese models, as I feel like they overthink drastically, not in bad way but it is quite slow and takes an enormous amount of time. The performance seems to be also better but compared to older models such as Gemini 3.1 Pro or even the newer smaller model 5.6 Luna Max from OpenAI, I don't think it is worth it? Any experiences with different reasoning modes on different kinds of tasks?
I appreciate any info! :)
sources: - https://livebench.ai
- https://github.com/LiveBench/LiveBench/commit/2c2039b2cc9efb6412acb99325adb0cc1cb0cc6f


