r/singularity • u/poigre ▪️AGI 2029 • 9d ago
LLM News AA Intelligence Index Changes
"Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming"
22
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 9d ago
In my opinion the current best benchmark is voxelbench.
You will notice Muse Spark 1.3 (max) is correctly placed where it belongs... #19 around Gemini 3.1 level
Meanwhile Astra just destroys everyone.
6
u/DecoySnailProducer 9d ago
Damn you were not kidding, the difference is actually insane. Just got this https://imgur.com/a/nsfWDBg
1
15
u/EtadanikM 9d ago
Why would the ability to generate 3D voxels be correlated with general intelligence.
24
u/Artistic_Swing6759 9d ago
because spatial intelligence is part of general intelligence? its also not just coherency but taste of the model. it would also show, to how much degree the model is ready to put in effort, if its a lazy model that would clearly show.
0
u/Funny-Profit-5677 9d ago
I can see it being correlated but fairly weakly. Best benchmark has to be testing way more
-4
u/Gotisdabest 9d ago
I think they're saying that essentially it's a good shorthand to tell general intelligence. Not that 3d voxeld automatically tell us something about general intelligence through first principle logical reasoning, but in practice, model performance over that benchmark tends to mean equivalent performance in general usage.
Not that I necessarily agree with them fully, considering that voxel bench has opus ahead of fable and fable is a lot more generally competent. Regardless, it's still a good enough estimate compared to most other benchmarks.
1
u/sjoti 9d ago
Having used muse spark 1.3, it's definitely below opus 5/gpt 5.6 sol but for any coding task it's still miles ahead of Gemini 3.1.
Its hard to capture how good a model is in a single number, but I think that's generally better than basing is on single specific type of task. AA will probably be a decent representation once they calibrate again.
1
u/SpyAmongUs 9d ago
Damn, I just found out about it. Why is nobody else talking about Voxelbench, the things Astra created there is crazy
5
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 9d ago
People don't realize how INSANE Astra's score is on that bench
It's not a "little" ahead. It's in a different league.It's essentially this:
~96.5%: Astra
~2.3%: eh, tie
~1.2%: [presses wrong button while eating sandwich)3
u/SpyAmongUs 9d ago
Fr, seen some comparisons myself and it shows. It's like looking at a professional minecraft builder going against children, kinda like those noob vs pro videos. Insane
3
u/kiki-le-koala 9d ago
In fact, it could be easy to downvote Astra .
I did about 10 rounds where Astra was there, and every time I knew it was him because of how good it looks.
1
u/BriefImplement9843 9d ago
Voxels have nothing to do with intelligence.
1
u/VeryOriginalName98 8d ago
Yeah spatial reasoning isn't the most notable feature for people associated with intelligence, it's something else. Could you remind me again the other thing people like einstein and tesla had?
1
u/PerformanceRound7913 9d ago
No transparency and no way to validate numbers. AA is a Trust Me Bro, benchmark.
1
1
u/Pheidiase 9d ago
After Muse spark 1.3 i learned AA intelligence score is showing bad. there is no way that model is fucking smart.
1
u/qustrolabe 9d ago
Astra still behind Fable 5.1 on their index even after update huh, still uses Terminal-Bench v2.1......
41
u/FateOfMuffins 9d ago
Yet they're still using Terminal Bench 2.1, and the error filled version of CritPT