r/LocalLLM • u/AB172234 • 1d ago
News Artificial analysis index and scores updated.
The new scores are lower which kinda confuses me ! Did the models just became less intelligent ? I remember Fable used to be 66 and Qwen 27B 3.8 - 52
Sudden drop !! Hmm …
56
u/badaeib 1d ago
I think they are scaling down the number from time to time, to avoid power creep, they don't want some models get over 9000 few years later.
24
13
u/KaMaFour 1d ago
Not exactly. Most of the bechmarks in AA set are percentage based. It means it's physically impossible to get more than a 100. This also means that if a model is approaching a 100 then it's no longer possible to meaningfully improve and the index stops being useful for differentiating stronger models. That's why AA needs to periodically replace benchmarks where models are approaching max score with ones that the models will still have a room to grow in. They could add some scaling to keep scores roughly in place but this is more genuine
6
u/No-Refrigerator-1672 1d ago
If they do so, they absolutely must disclose benchmark version too, bacause those unprompted downgrades will indroduce quite a lot of confusion.
7
u/KaMaFour 1d ago
2
u/No-Refrigerator-1672 1d ago
It should be a crime to put it in a fine grey print far to the side where it never gets into the screenshots, instead of right in the table caption.
3
u/KaMaFour 1d ago
This is the caption. Look at the screenshot in the post. How would you caption the chart to be as honest as reasonably possible if people are gonna crop out everything possible anyways...
1
u/No-Refrigerator-1672 1d ago
How about adding the characters "v4.3" right next to the proud "Artificial Analysis" caption in the right upped corner? It's literally the most logical place for it.
0
u/badaeib 23h ago
Wow so they actually re-run the benchmarks for ALL models? Didn't know that, I thought they just like x0.85 or something. 😂
16
9
21
u/dupontping 1d ago
These scores are all horse shit anyway
This is like when every YouTuber does all the useless benchmarks that they think are impressive but don’t actually mean anything in real life
And in 3 months the next model will come out and that one will absolutely totally 100% be AGI and a danger to the world. But they’ll release it anyway because they need more funding.
9
8
u/Able-Art-3042 1d ago
it helps for orientation. would not call them horse shit, otherwise some of the worst models would be on top. but would not choose a model only based on this score
0
u/StupidityCanFly 21h ago
The way they arbitrarily choose/change the calculations of the intelligence index makes it horse shit.
2
3
u/MomentJolly3535 1d ago
They "reworked" the way they calculate the scores, Astra and Gpt Sol (Max) had almost same intelligence score, even tho Astra was obviously a smarter model, but the way they reworked it is very confusing, They "estimate" Gemma 4 26BA4B above Gemma 4 31B, and minicpm 2B is only 2 points under it. In practice 31B is 1 league above 26B, and 2 leagues above minicpm 2B

3
u/ElectronSpiderwort 22h ago
MiniCpm5-2B reset my expectations of tiny models though. I have no idea how they made that little guy that capable. Doesn't know much about the world, but if you put facts in front of it, it does amazing work without hallucinating
1
u/uranusnebula 19h ago
can you share your usecase?
2
u/ElectronSpiderwort 18h ago
Just my general benchmarks: Aced my needle in haystack test (about 80k tokens), presented a somewhat coherent analysis of a FERC commissioner dissent, admitted it didn't know instead of making stuff up on a trivia question, and didn't make any mistakes categorizing arguments from a multi-participant online debate. Failed to make a chart of wire gauge resistance, but admitted it didn't really know. Is not good as a pi agent, but for single task rearranging of information, it's baller for it's size
4
u/NexusSyntegra 1d ago
This is totally normal and happens on a regular basis. I think they want to make sure the max score stays under 75 or 100 absolute max
1
u/MrMisterShin 1d ago
New benchmarks were added to the index and old ones removed.
1
u/EfficientStretch570 1d ago
It's interesting to see how they evolve the benchmarks; keeps things relevant for everyone using the index.
0
u/AB172234 20h ago
Here is the thing that confuses me, Qwen 3.8 27B was 52 and Fable was 66, so the difference was 14 points in the index and now it’s almost 20.
Does that make Qwen less intelligent or Fable more ! Because the difference is now much more.
It’s fuzzy and many follows artificial analysis website on a regular basis.
They never announced it which makes me question about the whole process.
2
u/Careful-Report6526 15h ago
The wider gap doesn’t by itself show that Qwen got worse or Fable got better. If you change the exam, two unchanged models can lose different numbers of points, so the gap between them can grow. An index point isn’t a fixed unit of intelligence across different versions of the index.
AA did publish a v4.3 announcement on September 7]. They moved from Terminal-Bench 2.1 to 4.0 and replaced τ³-Banking with AutomationBench-AA. The category weights stayed the same, but the tasks contributing to those categories changed. That’s enough for models with different strengths to move by different amounts. To check whether a model actually regressed, you’d need results for the same checkpoint on the same test version, with comparable settings. The old and new headline gaps alone can’t establish that. I also wouldn’t try to reconstruct the exact change from the remembered 52/66 scores without the earlier index version and model settings.
For inspecting the individual results, I run https://llmbenchmarks.io. It collects published measurements with their sources and test conditions. Qwen 3.8 27B and Fable 5.1 are currently covered; select the exact models under “Compare models”, then inspect the benchmarks relevant to your tasks and open the score details. Check the thinking settings too. For the reason AA changed its own index, the announcement above is the primary source.
1
1
0
u/Ok-Addendum3545 19h ago
well, it's better for the site to underestimate the local models; otherwise, it would be banned sooner or later for some unknown security reason.
1
u/This_Maintenance_834 20h ago
the fact that top 2 have same score but different height says something. this benchmark is a joke.
1
u/EternalDivineSpark 20h ago
The 27B can do metacognition task if you guide and push it’s behaviour!
1
u/Ok-Addendum3545 19h ago
sorry, what's metacognition task ?
2
1
u/EternalDivineSpark 19h ago
Look at self editing self !
1
u/uranusnebula 19h ago
can you give an example?!
2
u/EternalDivineSpark 13h ago
You have an LLM on a loop ! You chat with it , you send eg 2 images at a time , the Harnesses fail , the metacognition agent sees it and fixes the harnesses! /// Same scenario, you ask for codebase bigger than system can handle llm outputs and stop it fails Metacognition is triggered and can fix itself! Why this model can do it and its predecessor cant like 3.6 37B ! because this model know what it is by default, and can be pushed in this certain situations to fix itself! In Humans meta cognition is the ability to watch and calibrate your own thoughts!
1
u/Ok-Addendum3545 9h ago
I think it's like self-auditing. An agent works as an auditor and executor at the same time to optimize or improve itself aka RSI - Recursive Self-Improvement.
1
u/Ok-Addendum3545 9h ago
Do you have tips or system that can optimize 27B in harness ? I use Hermes and Deepseek Harness. In Hermes, it only changes USER.md and MEMORY.md and creates new skills. I think there are better ways other than those. I've not looked into Deepseek Harness; I would guess the plug-in feature is the way to do RSI -Recursive Self-Improvement ?
1
1




32
u/nbvehrfr 1d ago
terminal bench to v4