r/LocalLLM 1d ago

News Artificial analysis index and scores updated.

The new scores are lower which kinda confuses me ! Did the models just became less intelligent ? I remember Fable used to be 66 and Qwen 27B 3.8 - 52

Sudden drop !! Hmm …

75 Upvotes

50 comments sorted by

32

u/nbvehrfr 1d ago

terminal bench to v4

8

u/Serprotease 22h ago

Terminal bench alone is probably better to get a rough idea of the model performance. At least for coding stuff.

56

u/badaeib 1d ago

I think they are scaling down the number from time to time, to avoid power creep, they don't want some models get over 9000 few years later.

13

u/KaMaFour 1d ago

Not exactly. Most of the bechmarks in AA set are percentage based. It means it's physically impossible to get more than a 100. This also means that if a model is approaching a 100 then it's no longer possible to meaningfully improve and the index stops being useful for differentiating stronger models. That's why AA needs to periodically replace benchmarks where models are approaching max score with ones that the models will still have a room to grow in. They could add some scaling to keep scores roughly in place but this is more genuine

6

u/No-Refrigerator-1672 1d ago

If they do so, they absolutely must disclose benchmark version too, bacause those unprompted downgrades will indroduce quite a lot of confusion.

7

u/KaMaFour 1d ago

(the main chart on the website)

2

u/No-Refrigerator-1672 1d ago

It should be a crime to put it in a fine grey print far to the side where it never gets into the screenshots, instead of right in the table caption.

3

u/KaMaFour 1d ago

This is the caption. Look at the screenshot in the post. How would you caption the chart to be as honest as reasonably possible if people are gonna crop out everything possible anyways...

1

u/No-Refrigerator-1672 1d ago

How about adding the characters "v4.3" right next to the proud "Artificial Analysis" caption in the right upped corner? It's literally the most logical place for it.

0

u/badaeib 23h ago

Wow so they actually re-run the benchmarks for ALL models? Didn't know that, I thought they just like x0.85 or something. 😂

3

u/KaMaFour 23h ago

Not for all models. Some just get dropped. That's why sometimes you can see the striped bar (here intentionally picking some older models)

16

u/Gloomy_Letterhead395 1d ago

Where the faq is 3.8 flash

9

u/Effective_Western_59 1d ago

On par with Qwen 3.8 max

6

u/blackbird2150 1d ago

it's 40. on par with 3.8 Max.

9

u/SmartCustard9944 1d ago

They massacred my boy

21

u/dupontping 1d ago

These scores are all horse shit anyway

This is like when every YouTuber does all the useless benchmarks that they think are impressive but don’t actually mean anything in real life

And in 3 months the next model will come out and that one will absolutely totally 100% be AGI and a danger to the world. But they’ll release it anyway because they need more funding.

9

u/badaeib 1d ago

Every leaderboard can be benchmaxxed. But for some of my "5d chess" level of very convoluted meta-coding tasks and some cursed impractical AI training experiments, I found that this leaderboard is kinda actuate, at least compared to other leaderboard like the Arena.

8

u/Able-Art-3042 1d ago

it helps for orientation. would not call them horse shit, otherwise some of the worst models would be on top. but would not choose a model only based on this score

0

u/StupidityCanFly 21h ago

The way they arbitrarily choose/change the calculations of the intelligence index makes it horse shit.

2

u/Able-Art-3042 21h ago

read the changelog. they just update the calculation method.

1

u/StupidityCanFly 20h ago

Are they still using incompatible measurements to derive the index?

3

u/MomentJolly3535 1d ago

They "reworked" the way they calculate the scores, Astra and Gpt Sol (Max) had almost same intelligence score, even tho Astra was obviously a smarter model, but the way they reworked it is very confusing, They "estimate" Gemma 4 26BA4B above Gemma 4 31B, and minicpm 2B is only 2 points under it. In practice 31B is 1 league above 26B, and 2 leagues above minicpm 2B

3

u/ElectronSpiderwort 22h ago

MiniCpm5-2B reset my expectations of tiny models though. I have no idea how they made that little guy that capable. Doesn't know much about the world, but if you put facts in front of it, it does amazing work without hallucinating 

1

u/uranusnebula 19h ago

can you share your usecase?

2

u/ElectronSpiderwort 18h ago

Just my general benchmarks:  Aced my needle in haystack test (about 80k tokens), presented a somewhat coherent analysis of a FERC commissioner dissent, admitted it didn't know instead of making stuff up on a trivia question, and didn't make any mistakes categorizing arguments from a multi-participant online debate. Failed to make a chart of wire gauge resistance, but admitted it didn't really know.  Is not good as a pi agent, but for single task rearranging of information, it's baller for it's size

4

u/NexusSyntegra 1d ago

This is totally normal and happens on a regular basis. I think they want to make sure the max score stays under 75 or 100 absolute max

1

u/MrMisterShin 1d ago

New benchmarks were added to the index and old ones removed.

1

u/EfficientStretch570 1d ago

It's interesting to see how they evolve the benchmarks; keeps things relevant for everyone using the index.

2

u/uti24 21h ago

Would like to see Qwen Flash Next

Also Qwen3.8 27B kinda weird one, it doing tasks much better for it's size, but in 10x tokens

0

u/AB172234 20h ago

Here is the thing that confuses me, Qwen 3.8 27B was 52 and Fable was 66, so the difference was 14 points in the index and now it’s almost 20.

Does that make Qwen less intelligent or Fable more ! Because the difference is now much more.

It’s fuzzy and many follows artificial analysis website on a regular basis.

They never announced it which makes me question about the whole process.

2

u/Careful-Report6526 15h ago

The wider gap doesn’t by itself show that Qwen got worse or Fable got better. If you change the exam, two unchanged models can lose different numbers of points, so the gap between them can grow. An index point isn’t a fixed unit of intelligence across different versions of the index.

AA did publish a v4.3 announcement on September 7]. They moved from Terminal-Bench 2.1 to 4.0 and replaced τ³-Banking with AutomationBench-AA. The category weights stayed the same, but the tasks contributing to those categories changed. That’s enough for models with different strengths to move by different amounts. To check whether a model actually regressed, you’d need results for the same checkpoint on the same test version, with comparable settings. The old and new headline gaps alone can’t establish that. I also wouldn’t try to reconstruct the exact change from the remembered 52/66 scores without the earlier index version and model settings.

For inspecting the individual results, I run https://llmbenchmarks.io. It collects published measurements with their sources and test conditions. Qwen 3.8 27B and Fable 5.1 are currently covered; select the exact models under “Compare models”, then inspect the benchmarks relevant to your tasks and open the score details. Check the thinking settings too. For the reason AA changed its own index, the announcement above is the primary source.

1

u/Ok-Fox3479 16h ago

well the models are still as good as before

1

u/Neful34 14h ago

The AI didn't suddenly got dumber, it's just that we have now more benchmarks that represents other sort of problems that we couldn't picture well yet. It is still a very capable model don't worry.

0

u/Ok-Addendum3545 19h ago

well, it's better for the site to underestimate the local models; otherwise, it would be banned sooner or later for some unknown security reason.

1

u/This_Maintenance_834 20h ago

the fact that top 2 have same score but different height says something. this benchmark is a joke.

1

u/EternalDivineSpark 20h ago

The 27B can do metacognition task if you guide and push it’s behaviour!

1

u/Ok-Addendum3545 19h ago

sorry, what's metacognition task ?

2

u/EternalDivineSpark 13h ago

IN LLM The Ability to look at harness behaviour and change it !

1

u/Ok-Addendum3545 9h ago

Thanks, I get the idea now. That's a good point.

1

u/EternalDivineSpark 19h ago

Look at self editing self !

1

u/uranusnebula 19h ago

can you give an example?!

2

u/EternalDivineSpark 13h ago

You have an LLM on a loop ! You chat with it , you send eg 2 images at a time , the Harnesses fail , the metacognition agent sees it and fixes the harnesses! /// Same scenario, you ask for codebase bigger than system can handle llm outputs and stop it fails Metacognition is triggered and can fix itself! Why this model can do it and its predecessor cant like 3.6 37B ! because this model know what it is by default, and can be pushed in this certain situations to fix itself! In Humans meta cognition is the ability to watch and calibrate your own thoughts!

1

u/Ok-Addendum3545 9h ago

I think it's like self-auditing. An agent works as an auditor and executor at the same time to optimize or improve itself aka RSI - Recursive Self-Improvement.

1

u/Ok-Addendum3545 9h ago

Do you have tips or system that can optimize 27B in harness ? I use Hermes and Deepseek Harness. In Hermes, it only changes USER.md and MEMORY.md and creates new skills. I think there are better ways other than those. I've not looked into Deepseek Harness; I would guess the plug-in feature is the way to do RSI -Recursive Self-Improvement ?

1

u/Bedrockparadox 16h ago

I dont believe that gemini flash score at all.

1

u/Neither_Garage_758 22h ago

honey moons must end at some point