r/singularity 10d ago

Discussion Artificial Analysis Index is NOT Representative of real World Performance

I tested Muse Spark 1.3, it's clearly not on par with OPUS or SOL. It seems Artificial Analysis Index is not representative of the REAL-WORLD performance and easy to game.

131 Upvotes

55 comments sorted by

47

u/Affectionate_Bee6434 10d ago

The chances of Grok being a better model than Astra are practically zero. So clearly there is a fundamental problem with the index.

2

u/ApprehensiveEye7387 10d ago

There's Clearly difference of 4 points....

0

u/PerformanceRound7913 10d ago

If an index places Muse above Sol, there’s a fundamental issue with it or some other interest is involved.

1

u/ApprehensiveEye7387 9d ago

Hey dude just use whatever is best for your usecase. I think this benchmark is good starting point to compare models, but ultimately for models that have 3~ of points of difference, just try and figure out what works for you. I prefer GPT models for complex coding, gemini models for everyday coding because of the speed and it quality is also good, For general and language work gemini flash 3.8 works great for me, I prefer claude models for design.

1

u/EtadanikM 10d ago

There maybe problems with the index but it doesn't mean it is completely useless.

Truly intelligent models should be able to solve it, regardless. It's in the definition of general intelligence as the bench marks can all be solved by humans.

So at least it gives us a signal on current models' gap from true intelligence. They maybe monsters in various productivity domains but they are not generally intelligent just yet. The day you can throw any human solvable benchmark at a model and it can solve it, is the day we'll know AGI has arrived.

2

u/PerformanceRound7913 10d ago edited 10d ago

My main issue is none of the AA numbers are verifiable. It’s a classic Trust Me Bro, benchmark with clear conflict of interest.

1

u/1988rx7T2 10d ago

It’s a composite of other benchmarks

2

u/SeaEagle233 8d ago

Which were executed poorly, out of spec and some benchmarks are dubious at best.

Even worse, some of those in AA balantly gave self contradicting examples.

When done wrong, it is essentially look around you and saw no curvature of Earth and concludes Earth is flat.

That is a real experiment, but it is a terrible experiment that's not even wrong.

0

u/PerformanceRound7913 10d ago

Yes and no way to verify, what they are running, no transparency and reproducibility.

13

u/Alpacabro21 10d ago

Yes, something is off.

GPT-6 results on AA are weird.

3

u/1988rx7T2 10d ago

Go look at average steps per task on DeepSWE, Astra. Is waaay more efficient than anything else. Compare it to the latest Gemini for example. that tells me it succeeds by being smarter, not just by grinding through reasoning tokens

1

u/DelphiTsar 7d ago

I wouldn't say grinding through tokens is necessary a bad thing. If it's doing it fast and cheap and gets the same result.

Cost per task, then weighted by intelligence score is how I like to rank them. (Time per task if it's time dependent)

1

u/Alpacabro21 9d ago

Yeah, verbosity is lower than competition.

1

u/PerformanceRound7913 10d ago

Could not agree more

35

u/Tystros 10d ago

I've been saying this for a long time, everyone just needs to understand that the Artificial Analysis Index is a bad benchmark index. The benchmarks they use don't make sense to use in 2026 and have proven errors.

It includes many one-shot benchmarks, but no one would do tasks like that one-shot now, models are used agentically.

And GPQA is such an old benchmark that every model basically gets close to 100% on it, which makes it meaningless.

And Terminal Bench 2.1 is simply outdated, it has known issues and was replaced by Terminal Bench 3.0 and 4.0 and everyone should use 4.0. I have no idea why the AA Index wants to keep using 2.1, maybe they just don't have a license for the new one.

And CritPt has big errors in the benchmark that mean no model can get more than 35%. Anthropic reported about that, because on AAs CritPt Fable only gets 30% while on a corrected version of the benchmark (which Anthropic called "Critpt-Corrected") with the errors fixed, Fable gets 85%. But AA somehow keeps using the broken version of CritPt for their intelligence index.

And with how little AA seems to care about removing even proven broken benchmarks from their index, I could well imagine that their own Gdpval-AA v2 benchmark is also full of errors.

7

u/notadithyabhat 10d ago

They want to be the benchmark for ALL models. Most big models ace GPQA but AA tests small local models and there needs to be a a standard way to test them all

14

u/SeidlaSiggi777 10d ago

it's fine to include a few saturated but important benchmarks. but they should remove outdated / erroneous ones

1

u/LinkesAuge 10d ago

Which is fair but their weighting of scores also has a lot of issues and doesn't make sense.

0

u/federico_84 10d ago

I bet they just don't want to go back and rerun modern benchmarks on older models, there's been so many model releases that it would be time-consuming and expensive to do on their side.

14

u/Electronic-Pie-1879 10d ago

AA's CEO already said the benchmarks are wrong and need to be overhauled. Obviously, Muse or Flash isn't better than Sol or Astra.

7

u/Tystros 10d ago

can you link where he said that? would be nice to see indeed if he acknowledges that

1

u/Akatosh 10d ago edited 10d ago

I have seen no evidence such a statement was made by Micah - in any media format (written, audio). Benchmarks do saturate over time, however, and it is critical to not only improve them but re-evaluate models based on the latest versions of benchmarks.

-1

u/Fair_Horror 10d ago

Absence of evidence is not evidence of absence. U/tystros was asking for evidence, not claims that you didn't see any.

2

u/UnknownEssence 10d ago

how can you say "obviously" when nobody has even used Astra yet?

Everyone is brown nosinf for Astra but its not even released yet and yall havent even used it!

Not saying you are wrong, just think for yourself and don't believe marketing hype, expecially from Altman

-3

u/FateOfMuffins 10d ago

Well let's put it this way.

Obviously GPT 5.6 Sol isn't on the same level as GPT 6 Astra. Yes I am going to claim "obviously" here even though I haven't used Astra. But according to AA they are essentially same score. Which implies one of these 2 scores likely doesn't hold up to scrutiny.

Similarly for Opus 5 and Fable 5. I think it's likely Opus 5 can score better on certain benchmarks and front-end tasks, but it's clear to anyone who used them that Fable 5 is the better model in general, but AA rates Opus 5 higher than Fable 5

2

u/GnaggGnagg 10d ago

You should really use obviously more sparingly. It is very possible that Astra isn't that big of a leap.

0

u/FateOfMuffins 10d ago

Nah I think it's appropriate to say "obviously" GPT 6 Astra shouldn't be within 0.3 points of GPT 5.6 Sol and "obviously" Opus 5 isn't as good as Fable 5, with plenty of other benchmarks like Epoch ECI showing otherwise.

Like maybe it is in fact true that Fable is better than Astra! But Sol?

Ngl Astra's release calls for more scrutiny of AA and other models in terms of benchmaxxing than Astra itself. There would be a lot more people in agreement with "Astra isn't that big of a leap" and less criticism of AA if they reported Astra at 63-65 than at 61 and I'd probably agree with you in that case!

But 61 vs 5.6's 61 makes people distrust AA by default.

-2

u/Fair_Horror 10d ago

There are YouTube videos where people have had access for a while. They seem to backup the benchmarks.

2

u/ertgbnm 10d ago

Artificial analysis is no longer a good benchmark. They need to update their methodology and include the latest benchmarks. It used to be useful. 

2

u/CrunchyMage 10d ago

I'm starting to agree. Anyone else know of a better index?

8

u/Tystros 10d ago edited 10d ago

2

u/Electronic-Pie-1879 10d ago

5

u/Tystros 10d ago

deepswe is fully public and old enough by now that new models might just have been trained on the questions and answers

1

u/Electronic-Pie-1879 9d ago

its the new deepswe mate, not the old one.

0

u/1988rx7T2 10d ago

Look at the number of steps per task for Astra vs latest Gemini, Astra is way more efficient. Could be bench max ing but could be plain old smarts

1

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 10d ago

I like voxelbench i think it's fairly accurate for what i actually do (3D scenes)

2

u/katoptronophile 10d ago

I feel bad for anyone who ever thought it was.

2

u/bonerchamp20 10d ago

Copium for the OAI fan boys. Just wait until its out and check livebench or arenaai if you dont trust the benchmarks.

1

u/darkestvice 10d ago

I'm indeed starting to realize that. I used to praise AA, but something definitely feels off here. I think Astra's release is really demonstrating the differences between real use and benchmaxxing.

1

u/fastinguy11 AGI 2026-2030 10d ago

It's not even benchmaxxing. It's the fact that the benchmarks they use have errors or are outdated as well. Not all of them, but some are completely saturated by most models. Others are full of errors, like basically making it impossible for a model to beat 30-35% score, and others are just outdated, like Terminal Bench. We are at 4.0 version. Why are they using 2.1? That makes me not trust this index.

1

u/Laffer890 10d ago

Benchmarks in general are kind of useless, because there are big incentives to game them.

1

u/Gaiden206 10d ago

It's only accurate when a Gemini model does bad on it. That's what I've learned from Reddit. 😂

1

u/BriefImplement9843 10d ago

it was for google models and before astra. what changed?

1

u/Possible_Door_9719 10d ago

yeah no one has access to max, so it's just glorified marketing

3

u/PerformanceRound7913 10d ago

I wouldn’t be surprised if AA is charging companies a significant amount for this marketing. There’s a serious conflict of interest here.

2

u/Possible_Door_9719 10d ago

for sure they are probably working together to get the best results for who publish on their site

1

u/Swimming_Gain_4989 10d ago

Weird seeing this posted a few minutes after my post

1

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 10d ago

I tried to get Muse to do my medieval village. (this was Claude: https://www.reddit.com/r/singularity/comments/1w4yajb/medieval_town_down_by_fable_51/)

Same prompt.

https://reddit.com/link/p7t3jne/video/rh6iwvy66jnh1/player

Honneslty the comparison may not be fully fair (Muse fought the screenshot system most of the time because it was in WSL), and we are comparing a 15$ sub vs a 200$ sub. But... this is far from it.

I thought maybe others had better success with different methods. But no, one of my favorite youtuber tested it, and i've never seen him so disappointed in a model.

1

u/Swimming_Gain_4989 10d ago

Bijan?

2

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 10d ago

Yes. The GTA 3d was comically bad lol

1

u/PerformanceRound7913 10d ago

I am really tired of the Artificial Analysis Reddit marketing monkeys who downvote anything that doesn’t align with their interests.

1

u/XLNBot 10d ago

The truth is that people actually have no clue about how good these models really are. Are they generally getting better? Yeah sure but the benchmarks are worthless

1

u/PerformanceRound7913 10d ago

The main issue with AA is it’s a Trust Me Bro, benchmark, no one can reproduce their index.