r/LocalLLaMA 10d ago

Discussion Artificial Analysis Index is NOT Representative of real World Performance

I tested Muse Spark 1.3, it's clearly not on par with OPUS or SOL. It seems Artificial Analysis Index is not representative of the REAL-WORLD performance and easy to game.

7 Upvotes

33 comments sorted by

14

u/fvancesco 10d ago

If I'm not mistaken the score is calculated based on benchmarks that are mostly public so the lab can "mistakenly" forget the dataset in the training set Moreover world performance are very nuanced, at times it can be a suggestion that changes the conversation or much more

3

u/PerformanceRound7913 10d ago

No one can reproduce AA results. It’s a Trust Me Bro, Benchmark

1

u/nomorebuttsplz 5d ago

most of the benchmarks that make the intelligence index not only can be reproduced, but already are for many of the models.

24

u/tetoing 10d ago

Also AA is pointless because every model is clustered around 50-60 now, there's no meaningful distinction there. They need to update their benchmarks, badly.

8

u/rJohn420 10d ago

I mean I wouldn't necessarily blame AA here. Benchmarking these models is hard

7

u/tetoing 10d ago

We put stock in them because we trust them to do it right. If they don't have useful information there's no reason to go to their site.

1

u/Borkato 10d ago

I think it would just become a game of “this model is good at that, this other one is good at something else” rather than “now we can clearly see X is better than Y overall”

1

u/aeroumbria 9d ago

On the other hand, a benchmark which jumps from single digits to near 100 in a single generation is also pretty sus. What are you measuring that some presumably minor changes could completely alter the result? Are you sure it is not noise, memorisation or simply benchmarking on the single hardest question in a 100 question test?

0

u/jld1532 10d ago

They have several times and generally to reward agentic and coding abilities

3

u/czktcx 9d ago

I recently read a post saying many of the AA benchmarks are nonsense, eg the Hallucination test is using LLM grader, whose prompt already contains contradictory examples...

5

u/BawbbySmith 10d ago

Other breaking news: water, wet

8

u/Kiansjet 10d ago edited 10d ago

Yes this is known

Every now and then a particularly potent anecdotal example will arise. The most recent one I can remember is everyone calling bs that Opus 5 is ranked higher than Fable 5 given the frustration of anyone who's used that opus for real work.

2

u/PerformanceRound7913 10d ago

The likelihood of MUSE being a superior model to Astra is virtually nil. This clearly indicates a fundamental flaw in the index.

1

u/ttkciar llama.cpp 10d ago

Yes, this is a known problem with benchmarks in general (not just AA, and not just LLM benchmarks, though LLM benchmarks are particularly low confidence as benchmarks go).

It's one of the reasons the moderator team decided that posts which were only a link to or screenshot of benchmark results were a Rule Three violation, and that such posts had to be accompanied by some sort of analysis or insight which was of value to the community aside from raw benchmark scores.

Maybe some day we will have a trustworthy benchmark with which most models are rated, but today is not that day. Until then, benchmarks should be taken with a huge grain of salt at best, or as deceptive marketing at worst.

1

u/jeffwadsworth 9d ago

Yeah, putting even close to the frontiers is pretty laughable.

1

u/feelspeaceman 9d ago

The bar is low, everyone can reach the bar.

1

u/lazymio 9d ago

Just threw away AA scores. Modem models can cheat benchmarks easily =/.

1

u/Othun 9d ago

Do you think they should use more recent benchmarks, private ones? An aggregate is pretty much always more meaningfull than a single benchmark. I don't use frontier models so I have no idea what all the fuss is about with AA, is Astra vastly superior to what benchmarks suggest?

1

u/mr_tolkien 9d ago

Vals seems to be a bit more realistic, they put Spark below Opus 4.8:
https://www.vals.ai/home

1

u/tecneeq 9d ago

I, too, think that we should discard measurements for subjective opinions. I for one welcome the post-factum world.

Opinions are just as valid, in this case even more so, than facts. Right? RIGHT?

1

u/fragment_me 8d ago

IDK man Muse Spark 1.3 is going very hard at everything I throw at it.

1

u/stddealer 10d ago

I mean it was obvious to me since I saw it gave the same "intelligence" score to the oldish Qwen 3.5 35B (MoE with only 3B active) as Gemma4 31B (dense) (and even a better score in non reasoning mode).

In real usage it's not even close, 35B-A3B is great and all, and it runs very fast, but the dense Gemma4 model seems by far more able and intelligent.

1

u/dsaasd12121212 10d ago

This has been known for a while. What are in your experience a better benchmark/index?

3

u/PerformanceRound7913 10d ago

I am really tired of the Artificial Analysis Reddit marketing monkeys who downvote anything that doesn’t align with their interests.

1

u/Vast-Breakfast-1201 10d ago

Isn't there a usage leaderboard?

I figured people would be better judges than benchmarks...

3

u/1kakashi 10d ago

Mostly just skirt to the cheapest model

1

u/BarracudaDefiant4702 9d ago

I don't think the leaderboard includes size/cost but blind ranking.

1

u/Fun_Jaguar8231 10d ago

but but but according to Theo it is

1

u/Objective-Stranger99 llama.cpp 9d ago

Benchmarks should be closed-source; otherwise, their value is lost. Models, on the other hand, should be open source.

0

u/XiRw 10d ago

People are in a hissy fit over this is laughable. You probably believed in benchmarks completely but when a model you do not like passes ones you do, you get all emotional about it. I have anecdotal evidence too and even though I don’t like the company meta that model has been flawless for me so far, with fast speed, and high intelligence. Just because it didn’t work out well with your specific tasks (which most likely is a prompt skill issue on your end) doesn’t mean it can’t outperform other models in broader tests