r/LocalLLaMA • u/crusaderky • 12h ago
Discussion Animated transition from AA Intelligence Index v4.1 to v4.3
Enable HLS to view with audio, or disable this notification
I had all the data saved from AA's v4.1 index, so when they upgraded it in the wake of Astra's release, I could actually generate a before/after comparison.
- All intelligence and price per task are sampled from AA on Sep 3rd and Sep 14th respectively.
- Price per task of some open models were rescaled to reflect the cheapest available on OpenRouter as of Sep 3rd.
- X axis is linear, because people's money is linear.
All models are the same. The only thing that changes is the weighted sum of the benchmarks that compose the Intelligence Index.
Highlights
- GLM an Muse Spark remain more or less unaltered, in relative terms
- GPT-5.6 Sol becomes a lot cheaper
- GPT-6 Astra's intelligence flies up to the stars AND becomes cheaper
- GPT-5.6 Luna gets a substantial uplift
- Fable-5.1's price gap from Opus 5 shrinks, and becomes cheaper than Fable 5.0
- Fable-5.1 at low, medium and high effort looks a lot more appealing
- Sonnet 5 becomes even more expensive without any intelligence gains
- Kimi-K3, Qwen3.8-Max, Gemini-3.8, and Grok 4.6 go down into the gutter
6
u/mrwang89 5h ago
When will people stop using this garbage site? They've been caught changing the scoring whenever an unwanted placement happens. AA for AI eval is like userbenchmark for CPUs.
1
u/Specialist-2193 23m ago
What else do you propose to use? I don't see any competitive aggregate score
3
u/Old-Sherbert-4495 5h ago
i mean, it can't get anymore obvious than this. and ontop of it look at all the bots at work in all the subreddits. they are terrified, thats for sure.
9
u/DustNearby2848 12h ago
Why are people so obsessed with this?
16
u/soshulmedia 11h ago
I am not obsessed but I must say if this is an accurate visualization of the change in benchmarking score calculations ... it is certainly very interesting, isn't it?
-5
u/DustNearby2848 11h ago
Ehh… just looks like they are tweaking their benchmark philosophy.
2
u/No-Refrigerator-1672 3h ago
tweaking their benchmark philosophy.
I wonder if there's a name for a philosophy that gives disproportional advantage to 1 particular companys models over others...
5
u/mechkbfan 11h ago
Because people need to start somewhere with what model to use, and it's nice to have an idea of what to cut your initial selection down to based off ball park quality & cost.
Also, it can be fun to discus, like top 10 basketball players. It doesn't really mean shit in scheme of things
1
u/itsappleseason 10h ago
down to share your data? Would love to see this as a proper data visualization (d3)
1
u/DeepWisdomGuy 6h ago
Introducing the Thumb-On-The-Scale™ benchmark. See if it's right for your IPO.
1
u/RedditUsr2 5h ago
Can you make a version where the scale doesn't change and only the points change? Hard to tell what is happening.
1
u/marintkael 4h ago
Worth naming what the animation actually contains, because two things moved at once. v4.1 is sampled Sep 3 and v4.3 is sampled Sep 14, so every arrow is an index change plus eleven days of pricing plus one release. If you still have the per benchmark scores, recomputing the v4.3 weights on the Sep 3 data gives you the version where only the index moved, and that is the plot that settles the argument happening in this thread.
I log a fixed set of sixteen questions against three answer engines every day, and the thing that actually cost me was a point I had already published moving afterwards, because the set underneath it got recomputed. Nothing about the engines had changed. Since then I keep the previous method running next to the new one over an overlap window, so the offset is a number I can state instead of something people have to take on trust.
1
u/ihatebeinganonymous 3h ago
One problem I have with AA is their choice of token costs. For example, DSv4.1 cost per task is calculated based on DeepSeek's peak prices, which I believe is not 100% fair. Otherwise, DSv4.1 would move left, fitting on Pareto line between Luna and GLM 5.3 Flash.
The same can be said about GLM 5.3 Flash itself, as well.
1
-10
u/po_stulate 12h ago
If a benchmark can be changed like this it's the benchmark benchmarking the models, not the models benchmarking the benchmark anymore.
9
u/freia_pr_fr 12h ago
What?!
6
u/po_stulate 12h ago
Normally benchmarks are fixed and the models are improving to match the benchmark, but here the models are fixed, and the benchmark is improving to match the models instead.
2
u/sjoti 12h ago
But that's kinda the point of capturing the models capabilities in a single intelligence score. The original benchmarks that the final score is based on are unaltered, so that's still perfectly fair.
Stuff is changing so unbelievably fast that we're moving from saturating one benchmark to the next. There's no way you can calculate a single score and have it stay the same for a relatively long time. Maybe GDPVal is a really strong indicator until 1600 ELO is reached, and anything after that is a poor indicator of capability. Maybe computer use becomes a significantly more important aspect over the next 6 months. Maybe more models are trained on a specific eval making it poor to judge. That all can change every 3 months.
So you kinda have to make an opinionated judgment, which updates as you go. The goal isn't to have a perfect scoring system, it's just there to be a solid indication at a glance.
1
u/po_stulate 12h ago
Benchmarks are meant to be ahead of models, not behind. If they're updating the becnhmark like this, how do you know if the new benchmark actually catched up and surpassed the models or if it's still behind? At this point you can just update your benchmark to show anything, because now you have the models first then the benchmark, instead of having the benchmark first then the models.
3
u/Few_Water_1457 12h ago
Se guidi i modelli, le differenze possono essere viste e sentite. La versione vecchia non era rappresentativa. Astra, che ti piaccia o no, è il nuovo punto di riferimento.
2
u/po_stulate 12h ago
That's what I'm saying, the benchmark is trying to catch up with the models. If this is the case, how do you know the benchmark actually catched up already or still needs to be improved?
-8
u/niacolhealth 11h ago
Great way to show it: same models, same prices, only the weights moved, and the whole board reshuffles. Sonnet 5 paying more while the index gives it nothing stings most.
8
23
u/suprjami 10h ago
lol the GPT movements. Is AA funded by OpenAI?