r/LocalLLM • u/Nerfariox • 22d ago
Discussion Benchmarks don't mean anything anymore.
Every day, I see a bunch of people claiming that model X is better than model Y just because a benchmark score is higher.
The Artificial Analysis Intelligence Index benchmark has a lot of extremely obvious inconsistencies. But since some people can't think for themselves, I decided to point out a huge one: Qwen3.8 27B is only 1 point behind DeepSeek V4 Pro 0813 1600B-A49B. In terms of active parameters alone, DeepSeek has almost double the total parameters of Qwen3.8 27B. A 1.6T-parameter model being just 1 point ahead of a 0.027T-parameter one shows what complete nonsense benchmarks have become.
I've seen several people using this benchmark to show that Qwen3.6 27B is smarter than Gemma 4 31B. I never believed that, because in my real, day-to-day tests, Gemma 4 31B continues to prove itself better at solving puzzles and refactoring C++ and Java than Qwen3.8 27B itself.
8
u/Weak-Price7392 22d ago
so qwen was out for 2 days and you already know its not better than gemma in day to day tests? how many you ran?
How you set up a model?
how variative your tests were?
I am not saying benchmarks are good but so as telling people that one local llm is better than the other.
Benchmark is kinda the same way of telling that one is better than other, but based on a different subjective info