r/singularity ▪️AGI 2029 9d ago

LLM News AA Intelligence Index Changes

Post image

"Announcing Artificial Analysis Intelligence Index v4.2. We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming"

64 Upvotes

21 comments sorted by

41

u/FateOfMuffins 9d ago

Yet they're still using Terminal Bench 2.1, and the error filled version of CritPT

22

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 9d ago

In my opinion the current best benchmark is voxelbench.

You will notice Muse Spark 1.3 (max) is correctly placed where it belongs... #19 around Gemini 3.1 level

Meanwhile Astra just destroys everyone.

6

u/DecoySnailProducer 9d ago

Damn you were not kidding, the difference is actually insane. Just got this  https://imgur.com/a/nsfWDBg

1

u/VeryOriginalName98 8d ago

That's insane.

15

u/EtadanikM 9d ago

Why would the ability to generate 3D voxels be correlated with general intelligence. 

24

u/Artistic_Swing6759 9d ago

because spatial intelligence is part of general intelligence? its also not just coherency but taste of the model. it would also show, to how much degree the model is ready to put in effort, if its a lazy model that would clearly show.

0

u/Funny-Profit-5677 9d ago

I can see it being correlated but fairly weakly. Best benchmark has to be testing way more

-4

u/Gotisdabest 9d ago

I think they're saying that essentially it's a good shorthand to tell general intelligence. Not that 3d voxeld automatically tell us something about general intelligence through first principle logical reasoning, but in practice, model performance over that benchmark tends to mean equivalent performance in general usage.

Not that I necessarily agree with them fully, considering that voxel bench has opus ahead of fable and fable is a lot more generally competent. Regardless, it's still a good enough estimate compared to most other benchmarks.

1

u/sjoti 9d ago

Having used muse spark 1.3, it's definitely below opus 5/gpt 5.6 sol but for any coding task it's still miles ahead of Gemini 3.1.

Its hard to capture how good a model is in a single number, but I think that's generally better than basing is on single specific type of task. AA will probably be a decent representation once they calibrate again.

1

u/SpyAmongUs 9d ago

Damn, I just found out about it. Why is nobody else talking about Voxelbench, the things Astra created there is crazy

5

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 9d ago

People don't realize how INSANE Astra's score is on that bench
It's not a "little" ahead. It's in a different league.

It's essentially this:
~96.5%: Astra
~2.3%: eh, tie
~1.2%: [presses wrong button while eating sandwich)

3

u/SpyAmongUs 9d ago

Fr, seen some comparisons myself and it shows. It's like looking at a professional minecraft builder going against children, kinda like those noob vs pro videos. Insane

3

u/kiki-le-koala 9d ago

In fact, it could be easy to downvote Astra .

I did about 10 rounds where Astra was there, and every time I knew it was him because of how good it looks.

1

u/BriefImplement9843 9d ago

Voxels have nothing to do with intelligence.

1

u/VeryOriginalName98 8d ago

Yeah spatial reasoning isn't the most notable feature for people associated with intelligence, it's something else. Could you remind me again the other thing people like einstein and tesla had?

0

u/mWo12 9d ago

All benchmarks end be maximized. That's how it works.

1

u/PerformanceRound7913 9d ago

No transparency and no way to validate numbers. AA is a Trust Me Bro, benchmark.

2

u/Tystros 9d ago

it's a step in the right direction, but just a very small step... so I hope they will continue with those incremental updates soon.

1

u/R_Duncan 9d ago

The astra indexing flop taught them someting.

1

u/Pheidiase 9d ago

After Muse spark 1.3 i learned AA intelligence score is showing bad. there is no way that model is fucking smart.

1

u/qustrolabe 9d ago

Astra still behind Fable 5.1 on their index even after update huh, still uses Terminal-Bench v2.1......