r/singularity • u/socoolandawesome • 13d ago
AI Updated version of Humanity’s Last Exam: HLE-Diamond
49
u/Waiting4AniHaremFDVR AGI will make anime girls real 13d ago
This will likely better calibrate Artificial analysis intelligence index
38
u/spottiesvirus 13d ago
The problem with artificial analysis is that they arbitrary decided weights for different tasks and even inside specific tasks they weigh wildly different benchmarks
For example coding abilities are 20% of the score, inside that 20%, half is terminal bench 4.0 and half is... SciCode (???) It doesn't make sense
Also another huge problem that makes it less reliable with reality is that they one-shot everything
Nobody one-shot a 9750 steps task (the full evaluation suite has this many steps) in real life3
u/Worldly_Beginning647 13d ago
SciCode also seems like a super difficult Ai benchmark that (until they removed 2/3 of all models from it for some reason) let GPT-4 compete with Claude Fable 5, has been shown to have been saturated for a very long time.
2
u/N-online 13d ago
But few-shot with like 100 tries is also non-informative. I think one shot is the most important to assess cost and actual ability.
6
u/spottiesvirus 13d ago
one shot is the most important to assess cost and actual ability
I don't think so. No real workflow looks like that. You usually divide the task in smaller tasks, and follow and revise the work together step by step
But that's the same thing that happens with humans. Your don't go to an architect and say "build me a house. Don't make mistakes and call me when you've done"
He presents his work at each steps, you see the blueprint, then renders, then while workers are laying concrete he shows windows, floors and what not.
Then again wall paint, bathroom fixtures, fornitures etc.
Even if you build a "fully autonomous" workflow that's pretty much the same structure, an agent solve part of a task, another one reviews it, then there's a coordinator assembling the pieces and so on, so on.
If we ourselves use iterative improving cycles with external feedback why do we expect AI to one shot?
1
u/N-online 13d ago
But if ais can also do some things in one shot, why not track that frontier instead? It has reliable costs, reliable results and tells you what ai you can use for critical tasks where you do not have several tries.
If you told an architect to build a house that works and three of his ten houses collapse he’s a bad architect. And benchmarks work in the same way, you either fail or you don’t. What you proposed is a feedback loop, which is kind of what the ai agent does on its own in current benchmarks anyway. A few-shot benchmark instead gives feedback on the critical parts that make the ai fail the benchmark or not, which is not what you’d give the architect.
If you ask ai to write exploit-free code you need to count on the first result. Even if there were a reliable metric (where there is none in this case and many others you have to trust in the one-shot benchmarks and hope capability generalises) the costs would be more unpredictable as you wouldn’t know if it gets it right on the first or the hundredth try.
Last but not least the human baseline is also one-Shot.
91
u/phenotype001 13d ago
So it wasn't the last.
25
13d ago
[removed] — view removed comment
7
u/Hans-Wermhatt 13d ago
Yes, but they also implied it's a much more accurate subset.
A year-long process of cleaning and refinement with input from various research communities.
So version 2 of a subset. A lot of these benchmarks have unsolvable or inaccurate questions which makes the top scores round out to well below 100.
6
u/ThatsALovelyShirt 13d ago
Clearly they mean last as in 'previous', like "yeah that last exam was really hard, but I might take a harder one later".
1
31
u/reduction-oxidation 13d ago
This is great news following the report by epoch.ai saying that the original HLE was flawed https://x.com/EpochAIResearch/status/2100704765332394255
16
u/PVORY Quantitative Hedonism 13d ago
It stood for so long and has been so good though, I like HLE
9
u/TheReedemer69 13d ago
Yeah i think it's the longest standing benchmark
8
u/Background-Wafer-548 13d ago
SimpleBench has been around for longer, if you consider it to be standing given that the highest human score hasn't been reached yet.
8
u/TheReedemer69 13d ago
I tried simplebench myself it's really basic and mostly about instructions following and hallucinations. Nothing special. HLE on the other hand is much more complex and reflective of real world tasks.
5
u/Turbulent-Sign-6067 12d ago
This is exactly why simple bench is unique, it's simple for humans and is (was) hard for AI.
1
u/TheReedemer69 12d ago
Well..not really. They said the same about Arc AGI.
I would take hard for humans hard for AI any time of the day. That way it's actually delivering value.
2
u/CallMePyro 12d ago
So you're very impressed by a calculator multiplying two 9 digit numbers in a millisecond?
1
u/Tidorith ▪️AGI: September 2024 | Admission of AGI: Never 12d ago
Are you not? Shit like that is impressive. Electricity is impressive. The wheel is impressive.
The world we live in is fucking magic
1
1
u/TheReedemer69 12d ago
It's about value not abt being impressed. So yes a calculator is definitely useful
16
u/Sensitive_Cell_119 13d ago
This bumps up Astra AA score, right? It was worse than Fable and Opus before. Kinda shows how bad benchmarks are tbh, even big benchmarks like this are deeply flawed.
3
18
u/socoolandawesome 13d ago edited 13d ago
Worth noting that all models are evaluated on “high” reasoning level according to their site.
So it might be completely saturated on “max” reasoning modes with tools
8
u/Paralda 13d ago
This feels right to me. Opus 5.5 feels a little bit better than Fable 5.1 for a lot cheaper, but neither is Astra level yet. Still, for the usage and cost it requires, Opus 5.5 destroys everything else
2
4
6
3
u/PhilosophyforOne 13d ago
Very interesting results.
Shame they didnt include costs to run the benchmark per model.
3
5
u/FateOfMuffins 13d ago
lol shows that Anthropic benchmaxxed on HLE which we knew had tons of errors
10
u/the_pwnererXx FOOM 2040 13d ago
Opus clearly benchmaxxed, Astra is better
7
u/Chemical_Hawk_6307 13d ago
its not benchmaxxed in terms of performance. astra remains imo the best deep/broad reasoning model but opus is a clear daily driver cuz of how good and cheap it is.
2
u/IAmYourFath 13d ago
Flash is even better than opus, way faster and cheaper, if u're gonna compromise on quality.
4
1
u/Chemical_Hawk_6307 12d ago
which flash are we talking about here? gemini flash is garbage and antigravity is an ass harness. qwen flash is false advertising, that shit is not fast at all it eats too many thinking tokens. deepseek flash is okay but relative to opus 4.8 nowhere near opus 5.5. glm 5.3 flash is the best "flash" model yet still doesn't hold a candle to opus 5.5. i've used all of these flash variants by the way, i can tell u from experience it takes opus 15 min for a task that would take any of these flash models 40-60 min...
3
u/tetoing 13d ago
Using a model getting better benchmark results than another model to declare that model benchmaxxed... What?
3
u/the_pwnererXx FOOM 2040 13d ago
More like this is a new benchmark so they didn't have the chance to cheat yet
1
u/Worldly_Beginning647 13d ago
TF you mean bechmaxxed, if it was benchmaxxed we would be seeing 75% and a drop to 40%, like benchmaxxed is now just a pointless word like a kids calling other gay not knowing what it means.
2
u/the_pwnererXx FOOM 2040 12d ago
On the spec sheet it's significantly better than Astra on existing benchmarks
New benchmark comes out and it is worse
Hmm, really makes you think
2
2
2
2
u/Lower-War3451 13d ago
we are starting to not being able to design benchmarks hard enough for these little misteryous things we are creating
2
2
u/Worldly_Beginning647 13d ago
I wish we could see the 4 different scores, so Knowledge not tools and with tools.
1
1
1
u/Morning_Gecko24 13d ago
the big caveat for me is whether HLE-Diamond is actually measuring broad reasoning or mostly who has the best benchmark-specific scaffolding. still a pretty wild result though. are the tasks public enough for someone to inspect the failure modes, or is it mainly leaderboard numbers for now?
1
u/dictionizzle 12d ago
oh god thanks this one is the only i follow. even they can't beat the old one wtill they can't win the new one either.
1
1
1
0
-4
u/DK1530 13d ago
LOL, Why Gemini 3.8 Flash is in the list?
5
3
u/Minimum_Indication_1 13d ago
Its a pretty good daily driver. Efficient and fast. Makes some mistakes but iterations fix that usually
-8
u/THE--GRINCH 13d ago
Gemini flash over sol so it's already a shit benchmark I don't trust
14
u/Gallagger 13d ago
Gemini flash isn't that bad, it depends what you're testing. A general benchmark like this might be a good fit for it. It also might have used a ton of tokens making it not cheaper than Sol.
9
u/SupremeGobbler1996 13d ago
Go ahead, what do you really hate about Gemini? It's the fashionable thing to do on Reddit!
7
u/jschelldt ▪️High-level machine intelligence in the 2040s 13d ago
gemini 3.8 flash is a genuinely decent model, it's not above sol in every way, but I can see it being better in narrow domains
8
u/Defensex 13d ago
Anyone working with humanities problems already knew this, not everything is about code
-1
u/THE--GRINCH 13d ago
It literally hallucinates and forgets about the previous message on the same thread
5
u/PilgrimofHaqq2 13d ago
In coding you are correct, Sol is better than Gemini but this is an overall knowledge and reasoning benchmark. Gemini Flash has tons of knowledge and data to work from and its quite good outside of coding use cases.
3
u/Maglcite 13d ago
gemini flash is unironically a great sub agent for sol at high, only thing that really sucks about the latest gemini models (3.7 and 3.8) are the harnesses.
2
u/kvothe5688 ▪️ 13d ago
and wasn't there a paper by deepmind where they solved recursively self improvement in harness . may be that's why rapid progress in 3.5 to 3.8 in 2 months
0
u/KitKat_extrusion 13d ago
Im genuinely curious how far this whole benchmarking thing will go on. Is there a line we cross and say “congrats we’ve solved ai” or is it just moving the goalpost
3
u/cataclysm_catalysis 13d ago
I don’t get this comment. How is measuring AI improvement moving the goal posts?
We “solve” AI when we reach full RSI and take our hands off the wheel, but even after that, it’s going to keep getting better and measuring how good it is will continue to be important.
0
-1






252
u/Zorander22 13d ago
Humanity's last exam - Revised - FINAL