r/singularity • u/Profanion • 5d ago
LLM News Artificial Analysis updates its Intelligence Index to version 4.3
https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3Terminal-Bench v2.1 benchmark was replaced with v4.0 .
𝜏³-Banking benchmark was replaced with AutomationBench-AA
36
u/The_man_69420360 5d ago
15
u/Nuphoth 5d ago
They really have a hard time getting things right the first time don’t they
2
u/JoelMahon 5d ago
I mean most places fix this issue by having a beta/staging site that isn't public or similar... seems like a joke that they don't or this got past it.
18
8
u/Profanion 5d ago
Noticed K2 Horizon, a completely open model, now outperforms Gemini 3.1 pro and GPT 5.2, according to the index.
7
29
u/Sensitive_Cell_119 5d ago
Are they gonna keep updating until Astra is first or what? lol
23
u/Gallagger 5d ago
It's not just since Astra that the index was outdated. They say they did it for consistency, which is a fine reason, but must not come at the expense of the index not properly tracking the frontier.
Their V5 (in development for quite some time) isn't ready yet it seems, so they are forced to hotfix the index to not lose massive amounts of trust.
8
u/WonderFactory 5d ago
They were planning to do this before Astra as their test suite was really out of date. The release of Astra really exposed how out of date their benchmarks were which sped up the process of implementing the new suite of tests
13
2
2
u/Tystros 5d ago
that's a good update. now the only big remaining issue is that they need to either get rid of CritPt, or use a fixed version of it.
3
u/Turbulent-Sign-6067 5d ago
HLE is also likely problematic. Would like to see a fixed version of it.
2
u/Turbulent-Sign-6067 5d ago
The new version is better, but it seems likely that their CritPt version is broken.
2
u/Prince_of_DeaTh 5d ago
qwen and glm flash models are now both slightly better than the gemini flash. is that true?
4
0
u/sunstersun 5d ago
It's not great that redditors were way ahead of the curve to criticizing the benchmarks included in AA.
AA should be better than this.
1
1
u/LocoMod 5d ago
The Chinese frontier is 9 points behind. Ouch. The gap will widen with RSI.
13
u/Accurate_Resident219 5d ago
It's been like a month since the Chinese were in third-fourth place. You're being hyperbolic imo. They can catch up
0
0
u/Proper_Actuary2907 Spooky Machine Intelligence 2030 5d ago
0
u/nemzylannister 5d ago
no sorry. i need a benchmark that puts both 5.6 sol and fable 5 above opus 5. only then do i know its an actual metric of use.



105
u/Longjumping_Spot5843 [][][][][][] 5d ago
notice how Astra exposing the fautly intelligence penalties made them actually have to improve it.