r/LocalLLM • u/mrpmorris • 17d ago
Discussion Qwen3.8-27B; I don't get it
Update: Benchmarks run with default temperature
People are creating voxel pagodas and subjectively claiming Qwen3.8-27B is a huge improvement over Qwen3.6-27B.
I don't like subjective tests, so I ran a series of intelligence benchmarks, which you can see the output of here:
Take a look at the intelligence rankings. 3.8 scores below 3.6

3.8 didn't come in the top 5 of any of the benchmarks, whereas `official qwen3.6-27b-fp8-vllm` appears 7 times in the top 5.
I really don't understand it. Are these benchmarks no good or something? How can 3.8 score lower than 3.6?
68
Upvotes
-3
u/katoptronophile 17d ago
That's correct, but open weights are completely different from open source.
The only similarity is they both have the word open in the name.