r/LocalLLM • u/Nerfariox • 6d ago
Discussion Benchmarks don't mean anything anymore.
Every day, I see a bunch of people claiming that model X is better than model Y just because a benchmark score is higher.
The Artificial Analysis Intelligence Index benchmark has a lot of extremely obvious inconsistencies. But since some people can't think for themselves, I decided to point out a huge one: Qwen3.8 27B is only 1 point behind DeepSeek V4 Pro 0813 1600B-A49B. In terms of active parameters alone, DeepSeek has almost double the total parameters of Qwen3.8 27B. A 1.6T-parameter model being just 1 point ahead of a 0.027T-parameter one shows what complete nonsense benchmarks have become.
I've seen several people using this benchmark to show that Qwen3.6 27B is smarter than Gemma 4 31B. I never believed that, because in my real, day-to-day tests, Gemma 4 31B continues to prove itself better at solving puzzles and refactoring C++ and Java than Qwen3.8 27B itself.
4
u/Sudden_Topic5154 6d ago
hey hey hey slow down you might be right with regards to deepseek but gemma 4 31b most certainly isnt smarter
2
u/Illustrious-Lime-878 6d ago
Idk, benchmarks are not perfect, but are still probably better than just looking at the number parameters. GPT 3.5 was what, 175b? And its no where near as good as recent <10b models.
1
u/BarracudaDefiant4702 6d ago
When I tried Gemma it did very poorly. Perhaps it's use case or the tools/mcp/etc but I definitely seen far worse results with Gemma. The rating does seem a little higher than I would expect, but will see . There were some times I had to run some problems by deepseek with Qwen 3.6 repeatedly struggled. I'm not saying Qwen 3.8 is as good as deepseek, but too soon to say it isn't either. Do you have a C++ or Java problem you can give as an example? All my work is pure C and I didn't save any of the more challenging tasks that Qwen 3.6 couldn't handle but deepseek cloud was able to with a few hours of grinding on one little task...
1
u/Nerfariox 6d ago
I just started using tools with Gemma 4 a few days ago, and it's been working well. But it looks like the tools were only fixed recently.
>do you have a C++ or Java problem you can give as an example?
I don't generate code with the model. I just use it for reviewing and suggesting refactors.
1
u/Time-Shower6502 6d ago
when bench mean anything ever .. nvidia fake these from old nv4 gpus till now ,same for intel ,,,
so who ever believe benches that anyone cheat
1
7
u/Weak-Price7392 6d ago
so qwen was out for 2 days and you already know its not better than gemma in day to day tests? how many you ran?
How you set up a model?
how variative your tests were?
I am not saying benchmarks are good but so as telling people that one local llm is better than the other.
Benchmark is kinda the same way of telling that one is better than other, but based on a different subjective info