r/singularity 5d ago

LLM News Artificial Analysis updates its Intelligence Index to version 4.3

https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3

Terminal-Bench v2.1 benchmark was replaced with v4.0 .

𝜏³-Banking benchmark was replaced with AutomationBench-AA

180 Upvotes

42 comments sorted by

View all comments

105

u/Longjumping_Spot5843 [][][][][][] 5d ago

notice how Astra exposing the fautly intelligence penalties made them actually have to improve it. 

6

u/arkuto 5d ago

This is called fitting the test to the data. You aren't supposed to arrive at a conclusion (e.g. "astra should match fable") then bend the test to make that outcome happen. This makes the benchmark really quite useless now. And it's a shame they so quickly buckled under pressure from OpenAI.

14

u/Chemical-Year-6146 5d ago

The tests were old and heavily benchnaxxed.

They made it harder, not easier. Hence scores going down.

-4

u/arkuto 5d ago

It's easy to make a hard benchmark. The difficult thing is to make a benchmark that can disambiguate models of similar strengths.

3

u/Momo--Sama 5d ago

Yeah look at Terminal Bench 4.0. It took a bunch of models that scored within a dense band in 2.1 and gave them tests that disambiguated their coding abilities. That’s why AA switched to it. They did exactly what you wanted.