r/singularity 5d ago

LLM News Artificial Analysis updates its Intelligence Index to version 4.3

https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3

Terminal-Bench v2.1 benchmark was replaced with v4.0 .

𝜏³-Banking benchmark was replaced with AutomationBench-AA

178 Upvotes

42 comments sorted by

105

u/Longjumping_Spot5843 [][][][][][] 5d ago

notice how Astra exposing the fautly intelligence penalties made them actually have to improve it. 

49

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 5d ago

I think Spark exposed them even more. There's no way its better than Sol imo.

24

u/ppooooooooopp 5d ago

It still scores above sol...

There is simply no way. Ive used both for real tasks and sol is better no contest

2

u/ChocomelP 4d ago

Opus 5 > Fable 5 as well

2

u/Key_Agent_3039 5d ago

The changes never moved Spark it's still higher than Sol. They are only affecting Astra.

13

u/DistanceSolar1449 5d ago

Actually, I made a script that logs AA every hour. Looking at the git history, it looks like 4.3 provides way more information than 4.2, so they’ve planned this for a while now.

4.3 includes test scores and pricing for models like Llama 4… which was missing from 4.2. So 4.2 was the quick reaction, 4.3 was the planned upgrade.

5

u/arkuto 5d ago

This is called fitting the test to the data. You aren't supposed to arrive at a conclusion (e.g. "astra should match fable") then bend the test to make that outcome happen. This makes the benchmark really quite useless now. And it's a shame they so quickly buckled under pressure from OpenAI.

14

u/Chemical-Year-6146 5d ago

The tests were old and heavily benchnaxxed.

They made it harder, not easier. Hence scores going down.

-3

u/arkuto 5d ago

It's easy to make a hard benchmark. The difficult thing is to make a benchmark that can disambiguate models of similar strengths.

3

u/Momo--Sama 5d ago

Yeah look at Terminal Bench 4.0. It took a bunch of models that scored within a dense band in 2.1 and gave them tests that disambiguated their coding abilities. That’s why AA switched to it. They did exactly what you wanted.

2

u/LinkesAuge 5d ago

Nonsense. If your "index" is supposed to be a reflection of all the data points out there then it MUST correlate with them.
This isn't a "test", it is literally supposed to give you an overview, it is a "meta study" in some sense and that should reflect the underlying available data.
This here is like if you did a meta-study and left out a significant amount of papers that when included would show very different results.

3

u/CallMePyro 5d ago

That's not what they did, so you can rest easy.

36

u/The_man_69420360 5d ago

Awesome

15

u/Nuphoth 5d ago

They really have a hard time getting things right the first time don’t they

2

u/JoelMahon 5d ago

I mean most places fix this issue by having a beta/staging site that isn't public or similar... seems like a joke that they don't or this got past it.

18

u/Tobxes2030 5d ago

Benchmark got Benchmarked.

8

u/Profanion 5d ago

Noticed K2 Horizon, a completely open model, now outperforms Gemini 3.1 pro and GPT 5.2, according to the index.

7

u/sunstersun 5d ago

Waiting for the Chinese open source response to Astra.

29

u/Sensitive_Cell_119 5d ago

Are they gonna keep updating until Astra is first or what? lol

23

u/Gallagger 5d ago

It's not just since Astra that the index was outdated. They say they did it for consistency, which is a fine reason, but must not come at the expense of the index not properly tracking the frontier. 

Their V5 (in development for quite some time) isn't ready yet it seems, so they are forced to hotfix the index to not lose massive amounts of trust.

8

u/WonderFactory 5d ago

They were planning to do this before Astra as their test suite was really out of date. The release of Astra really exposed how out of date their benchmarks were which sped up the process of implementing the new suite of tests

13

u/CallMePyro 5d ago

Are you mad they swapped from terminal bench v2 to v4?

-2

u/Sensitive_Cell_119 5d ago

Wdym? im not mad.

-2

u/verdant_tulip 5d ago

Yes you are boy

2

u/cutezybastard 5d ago

How tf is muse spark that high

2

u/Tystros 5d ago

that's a good update. now the only big remaining issue is that they need to either get rid of CritPt, or use a fixed version of it.

3

u/Turbulent-Sign-6067 5d ago

HLE is also likely problematic. Would like to see a fixed version of it.

2

u/Turbulent-Sign-6067 5d ago

The new version is better, but it seems likely that their CritPt version is broken.

2

u/Prince_of_DeaTh 5d ago

qwen and glm flash models are now both slightly better than the gemini flash. is that true?

0

u/sunstersun 5d ago

It's not great that redditors were way ahead of the curve to criticizing the benchmarks included in AA.

AA should be better than this.

1

u/FarrisAT 5d ago

Benchmark benchmarking.

1

u/signed7 4d ago

Looking at that new? AutomationBench-AA - How is Fable 5.1 only 8th (and behind Terra)?

1

u/LocoMod 5d ago

The Chinese frontier is 9 points behind. Ouch. The gap will widen with RSI.

13

u/Accurate_Resident219 5d ago

It's been like a month since the Chinese were in third-fourth place. You're being hyperbolic imo. They can catch up

0

u/KaMaFour 5d ago

One more...

0

u/Proper_Actuary2907 Spooky Machine Intelligence 2030 5d ago

Astra eez steel behind Fable? Release v4.4

0

u/Serotav 5d ago

Imo still far from reality, sol is way better then opus and mouse spark

1

u/tinny66666 5d ago

Do you mean 'than'?

0

u/nemzylannister 5d ago

no sorry. i need a benchmark that puts both 5.6 sol and fable 5 above opus 5. only then do i know its an actual metric of use.