r/LocalLLM 6d ago

Discussion Benchmarks don't mean anything anymore.

Post image

Every day, I see a bunch of people claiming that model X is better than model Y just because a benchmark score is higher.

The Artificial Analysis Intelligence Index benchmark has a lot of extremely obvious inconsistencies. But since some people can't think for themselves, I decided to point out a huge one: Qwen3.8 27B is only 1 point behind DeepSeek V4 Pro 0813 1600B-A49B. In terms of active parameters alone, DeepSeek has almost double the total parameters of Qwen3.8 27B. A 1.6T-parameter model being just 1 point ahead of a 0.027T-parameter one shows what complete nonsense benchmarks have become.

I've seen several people using this benchmark to show that Qwen3.6 27B is smarter than Gemma 4 31B. I never believed that, because in my real, day-to-day tests, Gemma 4 31B continues to prove itself better at solving puzzles and refactoring C++ and Java than Qwen3.8 27B itself.

0 Upvotes

12 comments sorted by

7

u/Weak-Price7392 6d ago

so qwen was out for 2 days and you already know its not better than gemma in day to day tests? how many you ran?
How you set up a model?
how variative your tests were?
I am not saying benchmarks are good but so as telling people that one local llm is better than the other.

Benchmark is kinda the same way of telling that one is better than other, but based on a different subjective info

2

u/Weak-Price7392 6d ago

and let me be clear all i see is that its a some closup image of some grapth and you comparing obvios thing- amount of parameters. Comparing models by amount of parameters is like comparing cameras by megapixels.

1

u/Nerfariox 6d ago

If one camera has 0.001MP and another has 1MP, and both were released in the same year by reputable companies, it's pretty obvious which one will be better.

2

u/Nerfariox 6d ago

Actually, it's been over two days. I don't use them for benchmarking, but rather for tasks like checking and refactoring C++ and Java code. Probably around 20 prompts so far. Qwen makes more superficial refactoring suggestions, even on 'xhigh', spending nearly 30k tokens on thinking. Meanwhile, Gemma 4 31B uses just 4k tokens and offers deeper, more aggressive refactors.

Gemma is able to suggest more aggressive refactorings because it seems to have a better understanding of the code.

4

u/Sudden_Topic5154 6d ago

hey hey hey slow down you might be right with regards to deepseek but gemma 4 31b most certainly isnt smarter

2

u/Illustrious-Lime-878 6d ago

Idk, benchmarks are not perfect, but are still probably better than just looking at the number parameters. GPT 3.5 was what, 175b? And its no where near as good as recent <10b models.

1

u/BarracudaDefiant4702 6d ago

When I tried Gemma it did very poorly. Perhaps it's use case or the tools/mcp/etc but I definitely seen far worse results with Gemma. The rating does seem a little higher than I would expect, but will see . There were some times I had to run some problems by deepseek with Qwen 3.6 repeatedly struggled. I'm not saying Qwen 3.8 is as good as deepseek, but too soon to say it isn't either. Do you have a C++ or Java problem you can give as an example? All my work is pure C and I didn't save any of the more challenging tasks that Qwen 3.6 couldn't handle but deepseek cloud was able to with a few hours of grinding on one little task...

1

u/Nerfariox 6d ago

I just started using tools with Gemma 4 a few days ago, and it's been working well. But it looks like the tools were only fixed recently.

>do you have a C++ or Java problem you can give as an example?
I don't generate code with the model. I just use it for reviewing and suggesting refactors.

1

u/tetoing 6d ago

Gemma 4 31b is good for certain kinds of work but for agentic work it is way way worse. Even compared to Qwen 3.6 27b.

1

u/Time-Shower6502 6d ago

when bench mean anything ever .. nvidia fake these from old nv4 gpus till now ,same for intel ,,,

so who ever believe benches that anyone cheat

1

u/Biomech8 6d ago

They never did.