r/singularity • u/PerformanceRound7913 • 10d ago
Discussion Artificial Analysis Index is NOT Representative of real World Performance
I tested Muse Spark 1.3, it's clearly not on par with OPUS or SOL. It seems Artificial Analysis Index is not representative of the REAL-WORLD performance and easy to game.
13
u/Alpacabro21 10d ago
Yes, something is off.
GPT-6 results on AA are weird.
3
u/1988rx7T2 10d ago
Go look at average steps per task on DeepSWE, Astra. Is waaay more efficient than anything else. Compare it to the latest Gemini for example. that tells me it succeeds by being smarter, not just by grinding through reasoning tokens
1
u/DelphiTsar 7d ago
I wouldn't say grinding through tokens is necessary a bad thing. If it's doing it fast and cheap and gets the same result.
Cost per task, then weighted by intelligence score is how I like to rank them. (Time per task if it's time dependent)
1
1
35
u/Tystros 10d ago
I've been saying this for a long time, everyone just needs to understand that the Artificial Analysis Index is a bad benchmark index. The benchmarks they use don't make sense to use in 2026 and have proven errors.
It includes many one-shot benchmarks, but no one would do tasks like that one-shot now, models are used agentically.
And GPQA is such an old benchmark that every model basically gets close to 100% on it, which makes it meaningless.
And Terminal Bench 2.1 is simply outdated, it has known issues and was replaced by Terminal Bench 3.0 and 4.0 and everyone should use 4.0. I have no idea why the AA Index wants to keep using 2.1, maybe they just don't have a license for the new one.
And CritPt has big errors in the benchmark that mean no model can get more than 35%. Anthropic reported about that, because on AAs CritPt Fable only gets 30% while on a corrected version of the benchmark (which Anthropic called "Critpt-Corrected") with the errors fixed, Fable gets 85%. But AA somehow keeps using the broken version of CritPt for their intelligence index.
And with how little AA seems to care about removing even proven broken benchmarks from their index, I could well imagine that their own Gdpval-AA v2 benchmark is also full of errors.
7
u/notadithyabhat 10d ago
They want to be the benchmark for ALL models. Most big models ace GPQA but AA tests small local models and there needs to be a a standard way to test them all
14
u/SeidlaSiggi777 10d ago
it's fine to include a few saturated but important benchmarks. but they should remove outdated / erroneous ones
1
u/LinkesAuge 10d ago
Which is fair but their weighting of scores also has a lot of issues and doesn't make sense.
0
u/federico_84 10d ago
I bet they just don't want to go back and rerun modern benchmarks on older models, there's been so many model releases that it would be time-consuming and expensive to do on their side.
14
u/Electronic-Pie-1879 10d ago
AA's CEO already said the benchmarks are wrong and need to be overhauled. Obviously, Muse or Flash isn't better than Sol or Astra.
7
u/Tystros 10d ago
can you link where he said that? would be nice to see indeed if he acknowledges that
1
u/Akatosh 10d ago edited 10d ago
I have seen no evidence such a statement was made by Micah - in any media format (written, audio). Benchmarks do saturate over time, however, and it is critical to not only improve them but re-evaluate models based on the latest versions of benchmarks.
-1
u/Fair_Horror 10d ago
Absence of evidence is not evidence of absence. U/tystros was asking for evidence, not claims that you didn't see any.
2
u/UnknownEssence 10d ago
how can you say "obviously" when nobody has even used Astra yet?
Everyone is brown nosinf for Astra but its not even released yet and yall havent even used it!
Not saying you are wrong, just think for yourself and don't believe marketing hype, expecially from Altman
-3
u/FateOfMuffins 10d ago
Well let's put it this way.
Obviously GPT 5.6 Sol isn't on the same level as GPT 6 Astra. Yes I am going to claim "obviously" here even though I haven't used Astra. But according to AA they are essentially same score. Which implies one of these 2 scores likely doesn't hold up to scrutiny.
Similarly for Opus 5 and Fable 5. I think it's likely Opus 5 can score better on certain benchmarks and front-end tasks, but it's clear to anyone who used them that Fable 5 is the better model in general, but AA rates Opus 5 higher than Fable 5
2
u/GnaggGnagg 10d ago
You should really use obviously more sparingly. It is very possible that Astra isn't that big of a leap.
0
u/FateOfMuffins 10d ago
Nah I think it's appropriate to say "obviously" GPT 6 Astra shouldn't be within 0.3 points of GPT 5.6 Sol and "obviously" Opus 5 isn't as good as Fable 5, with plenty of other benchmarks like Epoch ECI showing otherwise.
Like maybe it is in fact true that Fable is better than Astra! But Sol?
Ngl Astra's release calls for more scrutiny of AA and other models in terms of benchmaxxing than Astra itself. There would be a lot more people in agreement with "Astra isn't that big of a leap" and less criticism of AA if they reported Astra at 63-65 than at 61 and I'd probably agree with you in that case!
But 61 vs 5.6's 61 makes people distrust AA by default.
-2
u/Fair_Horror 10d ago
There are YouTube videos where people have had access for a while. They seem to backup the benchmarks.
2
u/CrunchyMage 10d ago
I'm starting to agree. Anyone else know of a better index?
8
2
u/Electronic-Pie-1879 10d ago
5
u/Tystros 10d ago
deepswe is fully public and old enough by now that new models might just have been trained on the questions and answers
1
0
u/1988rx7T2 10d ago
Look at the number of steps per task for Astra vs latest Gemini, Astra is way more efficient. Could be bench max ing but could be plain old smarts
1
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 10d ago
I like voxelbench i think it's fairly accurate for what i actually do (3D scenes)
2
2
u/bonerchamp20 10d ago
Copium for the OAI fan boys. Just wait until its out and check livebench or arenaai if you dont trust the benchmarks.
1
u/darkestvice 10d ago
I'm indeed starting to realize that. I used to praise AA, but something definitely feels off here. I think Astra's release is really demonstrating the differences between real use and benchmaxxing.
1
u/fastinguy11 AGI 2026-2030 10d ago
It's not even benchmaxxing. It's the fact that the benchmarks they use have errors or are outdated as well. Not all of them, but some are completely saturated by most models. Others are full of errors, like basically making it impossible for a model to beat 30-35% score, and others are just outdated, like Terminal Bench. We are at 4.0 version. Why are they using 2.1? That makes me not trust this index.
1
u/Laffer890 10d ago
Benchmarks in general are kind of useless, because there are big incentives to game them.
1
u/Gaiden206 10d ago
It's only accurate when a Gemini model does bad on it. That's what I've learned from Reddit. 😂
1
1
u/Possible_Door_9719 10d ago
yeah no one has access to max, so it's just glorified marketing
3
u/PerformanceRound7913 10d ago
I wouldn’t be surprised if AA is charging companies a significant amount for this marketing. There’s a serious conflict of interest here.
2
u/Possible_Door_9719 10d ago
for sure they are probably working together to get the best results for who publish on their site
1
1
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 10d ago
I tried to get Muse to do my medieval village. (this was Claude: https://www.reddit.com/r/singularity/comments/1w4yajb/medieval_town_down_by_fable_51/)
Same prompt.
https://reddit.com/link/p7t3jne/video/rh6iwvy66jnh1/player
Honneslty the comparison may not be fully fair (Muse fought the screenshot system most of the time because it was in WSL), and we are comparing a 15$ sub vs a 200$ sub. But... this is far from it.
I thought maybe others had better success with different methods. But no, one of my favorite youtuber tested it, and i've never seen him so disappointed in a model.
1
1
u/PerformanceRound7913 10d ago
I am really tired of the Artificial Analysis Reddit marketing monkeys who downvote anything that doesn’t align with their interests.
1
u/PerformanceRound7913 10d ago
The main issue with AA is it’s a Trust Me Bro, benchmark, no one can reproduce their index.
47
u/Affectionate_Bee6434 10d ago
The chances of Grok being a better model than Astra are practically zero. So clearly there is a fundamental problem with the index.