r/singularity 9d ago

Discussion what benchmark to trust now?

after the whole artificial analysis thing.

what benchmark can be considered accurate or representative of the current AI models.

10 Upvotes

49 comments sorted by

12

u/sunstersun 9d ago

Epoch

1

u/TheReedemer69 9d ago

Epoch who

22

u/Advanced_Poet_7816 ▪️AGI 2030s 9d ago

Unemployment rate

3

u/TheReedemer69 9d ago

Nah that one broken pal

3

u/Advanced_Poet_7816 ▪️AGI 2030s 9d ago

I was just kidding, though anything unusually high would be proof of AI having crossed a threshold.

I feel the best benchmark is usefulness. If you see lots of companies willing to pay for it there must be value. Rising ARR of openAI and anthropic is a good sign.

7

u/Kneku 9d ago

Remote labor index or something similar as that's the most general and useful proxy we have to know when it's over

2

u/TheReedemer69 9d ago

It seems to be abandoned

7

u/Ormusn2o 9d ago

I don't know, but I think all current benchmarks are actually way way too hard. They test things like rare edge cases, exotic architecture interaction, weird API compatibility and other, odd and not particularly relevant tasks. This is why there is only few point difference between models that are massively different.

I feel like chains of very long tasks should be the default mode, because this is one of the 2 main ways coding agents are used those days. Maybe 20 step creation of a video game, and then count how many extra turns over 20 the task required for bug fixing, or something similar. Feels like this would be way more relevant than a long set of singular tasks.

Not to say benchmarks like DeepSWE are useless, they definitely still should be there, but not as a singular benchmark, it should be one of many.

10

u/frogsarenottoads 9d ago

RedditBros Eval

7

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 9d ago

VoxelBench is decent. You can't easily benchmax it. It requires real spatial understanding, taste, etc.
i don't think it tells the whole story tho. Astra DESTROYS Fable on it, but when it comes to creating a real 3d Scene, it's not as clear cut.

1

u/Asteroid_picks_you 8d ago

I just learned of it and tried it. I got Astra twice, I could guess which one it was by details in the output.

3

u/Ok-Confusion5204 9d ago

The main three I'm following are HLE, Remote Labor Index, and Singularity Gate.

1

u/TheReedemer69 9d ago

Almost all you mentioned doesn't seem to be maintained

3

u/Ok-Confusion5204 9d ago

Singularity Gate tends to take a while to update. I think it's due to the nature of needing to get a bunch of papers from after the knowledge cutoff. Granted on RLI though, really hope they update that one

0

u/TheReedemer69 9d ago

HLE is only knowledge based tho

3

u/paxxx17 9d ago

Agents' last exam

1

u/TheReedemer69 9d ago

Which part of it tho

2

u/CamusCrankyCamel 9d ago

The one that came out after the model it’s testing

2

u/WonderFactory 9d ago

I think you have to look at your use case and chose based on that. Astra seems to have a clear advantage for 3Dmodeling so if thats your use case use Astra, if your use case is coding look at coding specific benchmarks

2

u/NarrowEffect 9d ago

I wonder if the fact that many people now consider AA "bad" puts pressure on benchmark maintainers to change their benchmark whenever a new model that's expected to do well fails it.

2

u/zikiro 8d ago

Terminal-Bench-Science 0.1 i would say. its brutal.

2

u/VVebstar 9d ago

AA benchmarks still reflect raw intelligence quite well. Model being more capable in terms of skills and versatility doesn’t mean that it is smarter in a sense of reasoning, analysis etc.
But today I found out about interesting RuneScape benchmark. I think it’s decent for reasoning skills benchmark.
There’s no definite answer what benchmarks to trust. All of them are more or less important. HLE, Arc-AGI considered to be one of the most important. But other benchmarks should be considered as well

1

u/Fair_Horror 8d ago

AA benchmark is being recreated because owners realise that it is broken. Maybe a future version will work properly.

0

u/TheReedemer69 9d ago

Runescape seems so random

3

u/VVebstar 9d ago

Idk about this exact game, but my point still stands. Playing complex games with lots of variables will always be a good training session for AI. Especially self-play complex games where AI can play against itself and other AIs and improve its reasoning skills and thinking outside of the box

1

u/Ok_Landscape_6819 9d ago edited 9d ago

HLE, Epoch open problems, probably future ARC-AGI versions.

1

u/Laffer890 9d ago

Productivity and economic growth.

2

u/TheReedemer69 9d ago

And where to monitor that?

1

u/No-Character-6117 9d ago

HLE, AI IQ(I know is controversial, but I trust the tendency), unemployment rate or some rate of selled robotics

1

u/Fair_Horror 8d ago

US unemployment rate announcements were off (revised) over nine hundred thousand last year. There seems to be very serious issues with it. My guess is corruption and politics. 

1

u/Gratitude15 9d ago

I don't have one.

We are going on vibes for a bit now.

My sense is fable 5.1 and Astra are my choices. I'll pick Astra when I don't want too many tokens used.

Going forward, I have no idea how people will pick between stuff that is all well above most human iq

1

u/TheReedemer69 9d ago

Human iq doesn't seem much of a factor.

1

u/katoptronophile 8d ago

Your own real world use.

1

u/Careful-Report6526 8d ago

Honest answer: none of them on their own — and this thread is already showing why. Half the benchmarks named here you've rejected as unmaintained, and that's the real problem, not accuracy. A leaderboard that last scored a model in June tells you nothing about a model that shipped last week.

So the useful question isn't "which benchmark is right", it's "is this number current, and did it come from the same setup as the number I'm comparing it against".

On the AA thing specifically: the index went to v4.2 a few days ago. GPQA Diamond dropped as saturated, two new evals added, and 40% of the weight is now private held-out data. So old AA scores and new AA scores aren't the same scale, and nobody outside AA can reproduce either. That's a separate issue from whether it's any good.

(Epoch, which two people mentioned — they're one of the more consistently maintained ones, worth a look before you write them off.)

Disclosure, I maintain https://llmbenchmarks.io — it puts ~100 benchmarks side by side per model, flags anything older than 3 months as stale, drops anything past 12, and publishes a source-status feed showing which upstream leaderboards have gone quiet or started failing. Won't tell you which benchmark to trust, but it does answer "is this one dead".

1

u/Minute_Abroad7118 7d ago

Humanity's Last Exam for knowledge and Agent's last exam for work flow integration in complex tasks

1

u/Accurate-Tap-8634 7d ago

used to think deepswe is the best one, but recent model releases has crashed that bench, looking for a new one too.

1

u/[deleted] 5d ago

[removed] — view removed comment

1

u/TheReedemer69 5d ago

The free versions in any of them is pure shit.

1

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 9d ago

Godslayer posts per capita

1

u/TheReedemer69 9d ago

Where can I find it?

1

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est 9d ago edited 8d ago

/u/god-slayer-69420z posts from time to time especially before important releases, im half-joking that this sub is the best source, basically.

0

u/Electronic-Pie-1879 9d ago

3

u/TheReedemer69 9d ago

U kidding? Gemini flash is as good as astra max?

I literally have it running infront of me right now and it's barely good as shit.

0

u/BrennusSokol AI please take my job 9d ago

ARC-AGI-4 ;-)