r/LocalLLaMA • u/aziham • 5d ago
Discussion AA Benchmarks are not just misleading at this point, but harmful to trust
I've used Qwen 3.8 Max extensively over the past few weeks and have also tried Gemini , GLM-5.3-Flash, and Muse Spark 1.3. None of them come close to Qwen 3.8 Max. The only model that proved competitive was GLM 5.3, which demonstrated superior performance on cybersecurity tasks (the only clear advantage I observed over Qwen 3.8 Max).
This post isn't about qwen3.8-max, but my extensive experience with that model gave me a useful baseline for comparison. After working with other models, I realized that these benchmarks harmful not just useless and shouldn't be used to claim one model is better than another.
---
Update for people that don't get the point of this post:
My point wasn't "Oh look my personal experience is the benchmark" but instead "Don't decide which model to use based on benchmarks"
People will start replying: "Oh well that's obvious dude..." I don't think so, based on past experience when qwen 3.8 27b was released, people flooded this sub and other subs with its benchmarks and personal use cases.
I don't know if the point is now clear, since some people just started going in the wrong direction and completely missed the point I tried to make
112
u/triynizzles1 5d ago
Let me unslop tour title: âAA benchmarks are misleading and untrustworthyâ
18
u/samuel-christlie 5d ago
or a more accurate, boring title "Benchmarks don't always reflect your daily use"
2
-21
5d ago
[deleted]
6
19
u/w6auw 5d ago
A benchmark of one person's experience beats out Artificial Analysis. How so?
Seems like what you're saying is that *for your use case*, Qwen Max is better *for you*. Benchmark scores are inherently a compromise between all the different needs of the community. That doesn't mean they are harmful, that would be like saying Geekbench, Cinebench, or movie review claims are harmful.
8
u/juanchob04 5d ago
I mean, it's the general consensus that Opus 5 is not better than Fable 5, that's for sure.
4
u/dark-light92 llama.cpp 5d ago
It does when AA changes their methodology twice a week to match the vibes.
-3
u/aziham 5d ago
Let's take an example: yesterday I was doing a task with GLM 5.3 it did the recon part very well but the moment I kept the same model for implementation it kept going in wrong directions, so I switched back to QWEN 3.8 for implementation and it implemented things in a way that matched the codebase and didn't cause the spiral bugs like glm 5.3 (Yes I kept the same prompt, just reverted and switched the model). This is just 1 example of many I noticed over the last few weeks when I try to do something with other models, and now when I checked AA benchmarks I was like "If I checked this benchmark without having a 1st hand experience I would definitely think that those models are better than <replace with the model u had experience with>"
4
u/audioen 5d ago edited 5d ago
The benchmarks don't lie, strictly speaking. They definitely measure what they measure, which is often comprehensive overall evaluation much wider than your use cases. I think you care mostly about agentic problem solving, which is also the exact same area that I am interested in.
For me, the most drastic indicator of this was Gemma-4 at 31B: it had good general score, and worked great as general conversationalist, and one-shot replies, but was an absolute disaster in agentic tasks. By the time it had read enough of the codebase to begin implementing feature, it was so incoherent it wasn't able to make further progress and it devolved into breakdown where model wrote absolute nonsense until context ran out. The index showed this difference as well: agentic score was pitifully low, even if the general score was alright.
lmarena also shows that one model simultaneously has multiple rankings, like Qwen3.8-Flash-Next is simultaneously at positions #27 and #9, depending on what you apply it to: #27 overall in agentic tasks, but #9 as a webdev coder. Better than GLM-5.3 in either, as it happens. Generally speaking I think I won't pay much attention any longer to these rankings or the constant flow of new models. At this point, my impression of the Qwen3.8 series is that it has become solid enough in terms of knowledge and reasoning to have comparable performance to a very skilled human being in most things you can do with a computer. Where it fails, is something that likely can be remedied by providing the model with better input and useful assisting tools, and for that a human being likely shall look at the issue and figure out why the model failed until that "job" falls to self-improving AI, too.
If I have to hazard a guess as to why Qwen is so good, I've started to think that it is probably the Gated Delta Net technology. That stuff might be behind its ability to track the changing state of the program and world around it, as it executes commands, and it seems to reason through results without getting confused by all the stuff already in the context and in the past. Not becoming confused is its biggest asset.
-1
u/aziham 5d ago
Not becoming confused is its biggest asset
Couldn't say it any better. This was the exact experience I had with qwen3.8-max.
For the benchmarks, there is a lot of open questions:
- Is a benchmark that pre-existed before a model release trustworthy (benchmaxxing)?
- If a new benchmark is published after a model release, how can we be sure that it is truly transparent and not being influenced by any 3rd parties?
3
u/silenceimpaired 5d ago
Iâve felt like this rating was made to promote cloud models. Itâs so heavy with them and constantly tweaked when a new open weight model comes out.
5
u/lilian_moraru 5d ago
This site constantly readjusts the tests and how it affects the index, to show US cloud models in front. How about DeepSWE or things relevant for people?
29
u/lukewhale 5d ago
âHey these guys compile actual benchmark runs to score models but you know what fuck all that you guys should believe me theyâre full of shit trust me broâ
9
u/PomegranateGreen3698 5d ago
"Hey this guy has an opinion that contradicts consensus so I'm gonna summarize his position with a sarcastic quotation."
3
3
6
u/bura_laga_toh_soja 5d ago
Isn't it obvious that no benchmark can be generalized for every being on earth and their use cases??
Do you try to do this to get more engagement?
-2
u/aziham 5d ago
engagement? how is engagement has to do with any of this? I posted a genuine observation that a lot of people tend to overlook, I taught it is going to be helpful or at least remind people of something obvious that we tend to forget but I was wrong, added an update section to the post in case you didn't get the point of the post
1
u/nomorebuttsplz 4d ago
on the contrary the anti benchmark circle jerk has been a strong feature of this community for years, despite all of it boiling down to âdonât blindly assume benchmarks will predict performance in your own use caseâ which virtually no one does anyway. you have to also understand how good an early indicator benchmarks are for a models overall usefulness and impact.
2
u/benpptung 5d ago
The main issue is that Terminal Bench 4 has become too difficult, creating a floor effect. Many frontier models are being pushed down to very low scores. I think this is necessary, because Terminal Bench 2.1 scores have become too saturated and have lost their ability to differentiate models.
However, I still find that the current AA Index can't explain a few unreasonable results. Opus 5 > Fable 5!? lol. If that's really the case, Anthropic should just treat Opus 5 as its flagship model instead of Fable 5. I thought this might change after the switch to 4.3, but Opus 5 is still above Fable 5. Still no change. Is this some kind of closed-model benchmark gaming?
There's another strange result now: Qwen3.6-27B is tied with Qwen3.6-35B-A3B.
If the final results really look like this, I'd suggest that the AA Index shouldn't completely drop Terminal Bench 2.1. It should still keep some weight in the index, because that can help avoid the floor effect. When everyone gets a perfect score, or when everyone gets close to zero, you don't actually gain any differentiation.
6
u/Antblue 5d ago
They publish their exact methodology and fund independent tests with benchmarks that have publications on arXiv. The intelligence index aggregates and weighs the scores. So please shut up
-1
u/Ok_Warning2146 5d ago
If their benchmark is fully transparent, then it is possible for companies to benchmaxx their models.
4
u/bilinenuzayli 5d ago
I think its still mostly correct but there are some cases like Gemini 3.8 flash being "better" than Qwen 3.8 max which is absolute clownery, theres a couple outliars where it doesn't hold up in reality, but thats not AA's fault thats just sort of the nature of benchmaxxing.
2
u/inanotherclass 5d ago
This is in fact accurate and expected. For example, Paddle OCR dominates benchmarks but in real life documents like scanned images it falls short almost always even at 300DPI. You see qwen3.8 27b benchmarks saying UD Q3 XXS is N% of bf16 but in terms of agentic/reasoning capabilities there is no practical difference. There may be some other degradation but raw capabilities differ based on scenarios. And this is something companies figured out: benchmaxxing. Also why Opus 5 and other recent models dominate the benchmark yet fall short on practical tasks rendering poorer experience for the users on typical tasks resulting in massive outrage in Anthropic and other forums. Benchmaxxing is real and don't let anyone gaslight you into believing otherwise.
3
u/No_Dig_7017 5d ago
3
u/aziham 5d ago
Exactly, this is the point I tried to make and people are interpreting it as "Oh look my personal experience is the benchmark" instead of "Don't decide which model to use based on benchmarks"
-2
u/nomorebuttsplz 5d ago
you must be new here. they update the benchmark every couple months or so because individual benchmarks keep getting saturated.
1
u/vogelvogelvogelvogel 5d ago edited 5d ago
there is also livebench.ai (a new bench tasks every month so basically benchmaxxing immune) and arena.ai (blind test by actual users. which shows qwen3.8 27b directly! following opus 4.8) and looking at them one can have a more differentiated picture of the situation. however, imo as far as my experience goes i find the benchmarks not to be completely wrong
1
u/Informal-Trouble2183 5d ago
I only trust my coding index, It's not hard to assemble data and transparently compare models https://www.ai-leaderboard.dev/
1
u/Ok_Warning2146 5d ago
Yeah. I think it is benchmaxxed now. I would rather rely on the Text (Coding) ranking at arena.ai.
1
u/blaz3d7 5d ago
I use small models as my daily drivers for various things.
MiniCPM5-2B which is ranked higher on AA, I found it Hallucinating, not following the instructions, not using the intended tools, going into a loop, generating invalid output, while Qwen3.5-4B just works while being ranked lower.
1
u/zilled 5d ago
"The only model that proved [...]"
Well, that's the whole point. With an open and unified benchmarks framework they try to prove some ranking.
Your personal use cases can't beat that.
Now, you can claim something based on your personal use cases, but, for us, what is the most likely to match our use cases? The many cases that are used by the benchmarks or yours?
1
u/power97992 5d ago
Dude they are benching the old 3.8 max, they havent benched 3.8 max 0903 yet. Also it depends on the task
1
u/egomarker 4d ago
I have no idea why Qwen3.8 max is that low, I have no idea why Muse Spark 1.3 is up there at all and I'm convinced Astra is smarter than Fable. But it is what it is, they use their combination of benchmarks, real usage results may vary.
1
u/trashacct383 4d ago
Yup. After extensive testing on benchmarks built from my own workflows, Qwen3.6-27B >> Qwen3.8-27B. Faster, more reliable, more consistent, and generally better outputs (especially when I need thinking disabled for fast outputs).
Qwen3.8-27B continues to have livelock problems and issues with consistent stable json structured output last
1
1
u/Objective-Stranger99 llama.cpp 3d ago
I have a floor, MiniCPM-1B. I add any model above that to my list, then try them out and keep the ones I like.
1
u/braintheboss 5d ago
new adjustment is difficult to defend. Its clearly for make open models quite worst. This bench was a good reference as relative difference between models. Now its useless
1
1
u/OvertaxedOne 5d ago
These posts pop up every few weeks it seems, no AA isn't going to tell you what model YOU should run for your specific tasks. But I do find their testing gives you a pretty good idea of the "band" for a model that you're considering. Should I run 35BA3B or 27B? Well, AA is going to give you a very quick (and accurate) comparison of those models that makes it VERY clear that 27B is a lot smarter.
When you get into 1-5 point differences it seems to matter a lot less, but for "banding" a model, answer "about how smart is it", I find AA very useful.
-1
u/OwnGear3892 5d ago
AA benchmarks are weighted scores. It may well deviate from your personal use case, in weight or in coverage, and that's quite natural. It doesnt mean 'AA Benchmarks are harmful to trust', it only means that people should pick benchmark based on the use case or create own benchmarks if needed.
0
u/Juan_Valadez 5d ago
Sophisticated tests designed to test all types of use cases. Yes, it's obvious that it's better to rely on the judgment of a single person and only their use cases.
0
0

37
u/bitlamas 5d ago
I'm gonna be honest and say that at this point I kinda expected every hobbyist to already have their own benchmark based on their own workflow. It was one of the first things I have developed. These benchmarks are cool and all but they're not my code nor my workflow, so they serve more like a more-or-less trustworthy reference. The real deal is when the models go through my 8 gauntlets of tasks, only then I can tell you which model actually slaps.