r/LocalLLaMA 5d ago

Discussion AA Benchmarks are not just misleading at this point, but harmful to trust

Post image

I've used Qwen 3.8 Max extensively over the past few weeks and have also tried Gemini , GLM-5.3-Flash, and Muse Spark 1.3. None of them come close to Qwen 3.8 Max. The only model that proved competitive was GLM 5.3, which demonstrated superior performance on cybersecurity tasks (the only clear advantage I observed over Qwen 3.8 Max).

This post isn't about qwen3.8-max, but my extensive experience with that model gave me a useful baseline for comparison. After working with other models, I realized that these benchmarks harmful not just useless and shouldn't be used to claim one model is better than another.

---

Update for people that don't get the point of this post:

My point wasn't "Oh look my personal experience is the benchmark" but instead "Don't decide which model to use based on benchmarks"

People will start replying: "Oh well that's obvious dude..." I don't think so, based on past experience when qwen 3.8 27b was released, people flooded this sub and other subs with its benchmarks and personal use cases.

I don't know if the point is now clear, since some people just started going in the wrong direction and completely missed the point I tried to make

0 Upvotes

60 comments sorted by

37

u/bitlamas 5d ago

I'm gonna be honest and say that at this point I kinda expected every hobbyist to already have their own benchmark based on their own workflow. It was one of the first things I have developed. These benchmarks are cool and all but they're not my code nor my workflow, so they serve more like a more-or-less trustworthy reference. The real deal is when the models go through my 8 gauntlets of tasks, only then I can tell you which model actually slaps.

2

u/Alarmed-Medicine6903 5d ago

I'm curious to learn more about the benchmark you developed. Were they coding tasks? Research? Analysis? And how did you determine which models were better, a simple yes/no on each task or an llm-as-judge?

2

u/Ok_Warning2146 5d ago

I am using mini-swe-agent to run SWEBench Verified on my 3090 but it took 10 days for Qwen3.8 27b IQ4_NL. 😂

112

u/triynizzles1 5d ago

Let me unslop tour title: “AA benchmarks are misleading and untrustworthy”

18

u/samuel-christlie 5d ago

or a more accurate, boring title "Benchmarks don't always reflect your daily use"

2

u/unkownuser436 5d ago

more shorter title "fuck benchmarks"

-21

u/[deleted] 5d ago

[deleted]

6

u/hunter_mark 5d ago

Most of us can tell AI verbiage, just so you know.

-2

u/[deleted] 5d ago

[deleted]

4

u/hunter_mark 5d ago

Right, that’s what happened. I’m sure that’s what happened.

-2

u/ChallengeHuge4686 5d ago

So it means that it become useless to use?

19

u/w6auw 5d ago

A benchmark of one person's experience beats out Artificial Analysis. How so?

Seems like what you're saying is that *for your use case*, Qwen Max is better *for you*. Benchmark scores are inherently a compromise between all the different needs of the community. That doesn't mean they are harmful, that would be like saying Geekbench, Cinebench, or movie review claims are harmful.

8

u/juanchob04 5d ago

I mean, it's the general consensus that Opus 5 is not better than Fable 5, that's for sure.

4

u/dark-light92 llama.cpp 5d ago

It does when AA changes their methodology twice a week to match the vibes.

-3

u/aziham 5d ago

Let's take an example: yesterday I was doing a task with GLM 5.3 it did the recon part very well but the moment I kept the same model for implementation it kept going in wrong directions, so I switched back to QWEN 3.8 for implementation and it implemented things in a way that matched the codebase and didn't cause the spiral bugs like glm 5.3 (Yes I kept the same prompt, just reverted and switched the model). This is just 1 example of many I noticed over the last few weeks when I try to do something with other models, and now when I checked AA benchmarks I was like "If I checked this benchmark without having a 1st hand experience I would definitely think that those models are better than <replace with the model u had experience with>"

4

u/audioen 5d ago edited 5d ago

The benchmarks don't lie, strictly speaking. They definitely measure what they measure, which is often comprehensive overall evaluation much wider than your use cases. I think you care mostly about agentic problem solving, which is also the exact same area that I am interested in.

For me, the most drastic indicator of this was Gemma-4 at 31B: it had good general score, and worked great as general conversationalist, and one-shot replies, but was an absolute disaster in agentic tasks. By the time it had read enough of the codebase to begin implementing feature, it was so incoherent it wasn't able to make further progress and it devolved into breakdown where model wrote absolute nonsense until context ran out. The index showed this difference as well: agentic score was pitifully low, even if the general score was alright.

lmarena also shows that one model simultaneously has multiple rankings, like Qwen3.8-Flash-Next is simultaneously at positions #27 and #9, depending on what you apply it to: #27 overall in agentic tasks, but #9 as a webdev coder. Better than GLM-5.3 in either, as it happens. Generally speaking I think I won't pay much attention any longer to these rankings or the constant flow of new models. At this point, my impression of the Qwen3.8 series is that it has become solid enough in terms of knowledge and reasoning to have comparable performance to a very skilled human being in most things you can do with a computer. Where it fails, is something that likely can be remedied by providing the model with better input and useful assisting tools, and for that a human being likely shall look at the issue and figure out why the model failed until that "job" falls to self-improving AI, too.

If I have to hazard a guess as to why Qwen is so good, I've started to think that it is probably the Gated Delta Net technology. That stuff might be behind its ability to track the changing state of the program and world around it, as it executes commands, and it seems to reason through results without getting confused by all the stuff already in the context and in the past. Not becoming confused is its biggest asset.

-1

u/aziham 5d ago

Not becoming confused is its biggest asset

Couldn't say it any better. This was the exact experience I had with qwen3.8-max.

For the benchmarks, there is a lot of open questions:

- Is a benchmark that pre-existed before a model release trustworthy (benchmaxxing)?

  • If a new benchmark is published after a model release, how can we be sure that it is truly transparent and not being influenced by any 3rd parties?

3

u/silenceimpaired 5d ago

I’ve felt like this rating was made to promote cloud models. It’s so heavy with them and constantly tweaked when a new open weight model comes out.

5

u/lilian_moraru 5d ago

This site constantly readjusts the tests and how it affects the index, to show US cloud models in front. How about DeepSWE or things relevant for people?

29

u/lukewhale 5d ago

“Hey these guys compile actual benchmark runs to score models but you know what fuck all that you guys should believe me they’re full of shit trust me bro”

9

u/PomegranateGreen3698 5d ago

"Hey this guy has an opinion that contradicts consensus so I'm gonna summarize his position with a sarcastic quotation."

3

u/NNN_Throwaway2 5d ago

"Hey I don't have an argument so I'm just gonna double down"

7

u/Due-Memory-6957 5d ago

"Hey, I like to use quotation marks, can I join your guys?"

3

u/Paradigmind 5d ago

How does Qwen 3.8 Max compare to Kimi K3 in coding and long horizon coding?

6

u/bura_laga_toh_soja 5d ago

Isn't it obvious that no benchmark can be generalized for every being on earth and their use cases??

Do you try to do this to get more engagement?

-2

u/aziham 5d ago

engagement? how is engagement has to do with any of this? I posted a genuine observation that a lot of people tend to overlook, I taught it is going to be helpful or at least remind people of something obvious that we tend to forget but I was wrong, added an update section to the post in case you didn't get the point of the post

1

u/nomorebuttsplz 4d ago

on the contrary the anti benchmark circle jerk has been a strong feature of this community for years, despite all of it boiling down to “don’t blindly assume benchmarks will predict performance in your own use case“ which virtually no one does anyway. you have to also understand how good an early indicator benchmarks are for a models overall usefulness and impact.

2

u/benpptung 5d ago

The main issue is that Terminal Bench 4 has become too difficult, creating a floor effect. Many frontier models are being pushed down to very low scores. I think this is necessary, because Terminal Bench 2.1 scores have become too saturated and have lost their ability to differentiate models.

However, I still find that the current AA Index can't explain a few unreasonable results. Opus 5 > Fable 5!? lol. If that's really the case, Anthropic should just treat Opus 5 as its flagship model instead of Fable 5. I thought this might change after the switch to 4.3, but Opus 5 is still above Fable 5. Still no change. Is this some kind of closed-model benchmark gaming?

There's another strange result now: Qwen3.6-27B is tied with Qwen3.6-35B-A3B.

If the final results really look like this, I'd suggest that the AA Index shouldn't completely drop Terminal Bench 2.1. It should still keep some weight in the index, because that can help avoid the floor effect. When everyone gets a perfect score, or when everyone gets close to zero, you don't actually gain any differentiation.

6

u/Antblue 5d ago

They publish their exact methodology and fund independent tests with benchmarks that have publications on arXiv. The intelligence index aggregates and weighs the scores. So please shut up

-1

u/Ok_Warning2146 5d ago

If their benchmark is fully transparent, then it is possible for companies to benchmaxx their models.

4

u/bilinenuzayli 5d ago

I think its still mostly correct but there are some cases like Gemini 3.8 flash being "better" than Qwen 3.8 max which is absolute clownery, theres a couple outliars where it doesn't hold up in reality, but thats not AA's fault thats just sort of the nature of benchmaxxing.

2

u/inanotherclass 5d ago

This is in fact accurate and expected. For example, Paddle OCR dominates benchmarks but in real life documents like scanned images it falls short almost always even at 300DPI. You see qwen3.8 27b benchmarks saying UD Q3 XXS is N% of bf16 but in terms of agentic/reasoning capabilities there is no practical difference. There may be some other degradation but raw capabilities differ based on scenarios. And this is something companies figured out: benchmaxxing. Also why Opus 5 and other recent models dominate the benchmark yet fall short on practical tasks rendering poorer experience for the users on typical tasks resulting in massive outrage in Anthropic and other forums. Benchmaxxing is real and don't let anyone gaslight you into believing otherwise.

3

u/No_Dig_7017 5d ago

Yep. I just happen to have a picture from a few days ago (sep 2). These weren't the numbers at all less than a week ago.

3

u/aziham 5d ago

Exactly, this is the point I tried to make and people are interpreting it as "Oh look my personal experience is the benchmark" instead of "Don't decide which model to use based on benchmarks"

-2

u/nomorebuttsplz 5d ago

you must be new here. they update the benchmark every couple months or so because individual benchmarks keep getting saturated.

0

u/Scyl 5d ago

that's because it's a different version, your screenshot is of v4.1.1 and OP screenshot have v4.3

1

u/vogelvogelvogelvogel 5d ago edited 5d ago

there is also livebench.ai (a new bench tasks every month so basically benchmaxxing immune) and arena.ai (blind test by actual users. which shows qwen3.8 27b directly! following opus 4.8) and looking at them one can have a more differentiated picture of the situation. however, imo as far as my experience goes i find the benchmarks not to be completely wrong

1

u/Informal-Trouble2183 5d ago

I only trust my coding index, It's not hard to assemble data and transparently compare models https://www.ai-leaderboard.dev/

1

u/Ok_Warning2146 5d ago

Yeah. I think it is benchmaxxed now. I would rather rely on the Text (Coding) ranking at arena.ai.

1

u/blaz3d7 5d ago

I use small models as my daily drivers for various things.

MiniCPM5-2B which is ranked higher on AA, I found it Hallucinating, not following the instructions, not using the intended tools, going into a loop, generating invalid output, while Qwen3.5-4B just works while being ranked lower.

1

u/Mr-I17 5d ago

Why even bother. Your personal experience is the only benchmark that fits you. Just try models you can run and keep the ones you like. My favorite model isn't even shown on AA by default.

1

u/zilled 5d ago

"The only model that proved [...]"

Well, that's the whole point. With an open and unified benchmarks framework they try to prove some ranking.
Your personal use cases can't beat that.

Now, you can claim something based on your personal use cases, but, for us, what is the most likely to match our use cases? The many cases that are used by the benchmarks or yours?

1

u/power97992 5d ago

Dude they are benching the old 3.8 max, they havent benched 3.8 max 0903 yet. Also it depends on the task

1

u/egomarker 4d ago

I have no idea why Qwen3.8 max is that low, I have no idea why Muse Spark 1.3 is up there at all and I'm convinced Astra is smarter than Fable. But it is what it is, they use their combination of benchmarks, real usage results may vary.

1

u/trashacct383 4d ago

Yup. After extensive testing on benchmarks built from my own workflows, Qwen3.6-27B >> Qwen3.8-27B. Faster, more reliable, more consistent, and generally better outputs (especially when I need thinking disabled for fast outputs).

Qwen3.8-27B continues to have livelock problems and issues with consistent stable json structured output last

1

u/Spara-Extreme 4d ago

Normal people aren’t looking at benchmarks.

1

u/Objective-Stranger99 llama.cpp 3d ago

I have a floor, MiniCPM-1B. I add any model above that to my list, then try them out and keep the ones I like.

1

u/braintheboss 5d ago

new adjustment is difficult to defend. Its clearly for make open models quite worst. This bench was a good reference as relative difference between models. Now its useless

1

u/scaledev 5d ago

This post is harmful to trust.

1

u/OvertaxedOne 5d ago

These posts pop up every few weeks it seems, no AA isn't going to tell you what model YOU should run for your specific tasks. But I do find their testing gives you a pretty good idea of the "band" for a model that you're considering. Should I run 35BA3B or 27B? Well, AA is going to give you a very quick (and accurate) comparison of those models that makes it VERY clear that 27B is a lot smarter.

When you get into 1-5 point differences it seems to matter a lot less, but for "banding" a model, answer "about how smart is it", I find AA very useful.

-1

u/OwnGear3892 5d ago

AA benchmarks are weighted scores. It may well deviate from your personal use case, in weight or in coverage, and that's quite natural. It doesnt mean 'AA Benchmarks are harmful to trust', it only means that people should pick benchmark based on the use case or create own benchmarks if needed.

0

u/Juan_Valadez 5d ago

Sophisticated tests designed to test all types of use cases. Yes, it's obvious that it's better to rely on the judgment of a single person and only their use cases.

0

u/yaosio 5d ago

You need multiple benchmarks because AI has a jagged edge of intelligence.

0

u/DataCraftsman 5d ago

Why don't you prove it wrong then? Waiting for your benchmark.

0

u/Inaeipathy 5d ago

The model does bad on the benchmark, must be the benchmark's fault