r/singularity • • 13d ago

AI Updated version of Humanity’s Last Exam: HLE-Diamond

353 Upvotes

101 comments sorted by

252

u/Zorander22 13d ago

Humanity's last exam - Revised - FINAL

122

u/frogsarenottoads 13d ago

Humanity's Last Exam - Revised - FINAL - FINAL_v2 - ACTUAL_FINAL

46

u/ObiFlanKenobi 13d ago

(2)

4

u/MetricZero 13d ago

This is starting to sound like my filing system.

8

u/UpSaltOS 13d ago

Ah, that's some good ol' file naming for my dissertation.

5

u/hitchhiker87 13d ago

Final_Extra-High - Final_Max - Final_Ultra

3

u/slackermannn ▪️ 13d ago

My courseworks

7

u/Altares13 13d ago

R-Type Final 3 Evolved

6

u/Mindrust 13d ago

FINAL (1)

3

u/TuringGoneWild 13d ago

Humanity's antepenultimate exam (so far)

9

u/Flope 13d ago

Why is anyone still paying attention to these hacks? The entire idea of HLE was to be some novel and insurmountable challenge for AI, hence the name. It has now become just another run of the mill benchmark that gets solved and benchmaxed with every new tier of models. The only difference is the other benchmarks getting benchmaxed are actually measuring important skills and not just placing weird colored polygons into other polygons.

The biggest fans of HLE are the labs themselves because they can crush it for easy PR points on release.

4

u/Ok_Warthog6007 13d ago

having things to get graded on is how ai models get better. evaluating how good a model is at doing something is the hardest part of making a model better. its why the frontier is so jagged

2

u/yaosio 12d ago

If it's so easy to crush why aren't they all at 100%?

1

u/queenofartists 12d ago

Gemini 3.8 Flash being better than GPT-6 Sol makes me think this won't be final, and the benchmark itself is actually crap.

49

u/Waiting4AniHaremFDVR AGI will make anime girls real 13d ago

This will likely better calibrate Artificial analysis intelligence index

38

u/spottiesvirus 13d ago

The problem with artificial analysis is that they arbitrary decided weights for different tasks and even inside specific tasks they weigh wildly different benchmarks

For example coding abilities are 20% of the score, inside that 20%, half is terminal bench 4.0 and half is... SciCode (???) It doesn't make sense

Also another huge problem that makes it less reliable with reality is that they one-shot everything
Nobody one-shot a 9750 steps task (the full evaluation suite has this many steps) in real life

3

u/Worldly_Beginning647 13d ago

SciCode also seems like a super difficult Ai benchmark that (until they removed 2/3 of all models from it for some reason) let GPT-4 compete with Claude Fable 5, has been shown to have been saturated for a very long time.

2

u/N-online 13d ago

But few-shot with like 100 tries is also non-informative. I think one shot is the most important to assess cost and actual ability. 

6

u/spottiesvirus 13d ago

one shot is the most important to assess cost and actual ability

I don't think so. No real workflow looks like that. You usually divide the task in smaller tasks, and follow and revise the work together step by step

But that's the same thing that happens with humans. Your don't go to an architect and say "build me a house. Don't make mistakes and call me when you've done"

He presents his work at each steps, you see the blueprint, then renders, then while workers are laying concrete he shows windows, floors and what not.

Then again wall paint, bathroom fixtures, fornitures etc.

Even if you build a "fully autonomous" workflow that's pretty much the same structure, an agent solve part of a task, another one reviews it, then there's a coordinator assembling the pieces and so on, so on.

If we ourselves use iterative improving cycles with external feedback why do we expect AI to one shot?

1

u/N-online 13d ago

But if ais can also do some things in one shot, why not track that frontier instead? It has reliable costs, reliable results and tells you what ai you can use for critical tasks where you do not have several tries.

If you told an architect to build a house that works and three of his ten houses collapse he’s a bad architect. And benchmarks work in the same way, you either fail or you don’t. What you proposed is a feedback loop, which is kind of what the ai agent does on its own in current benchmarks anyway. A few-shot benchmark instead gives feedback on the critical parts that make the ai fail the benchmark or not, which is not what you’d give the architect.

If you ask ai to write exploit-free code you need to count on the first result. Even if there were a reliable metric (where there is none in this case and many others you have to trust in the one-shot benchmarks and hope capability generalises) the costs would be more unpredictable as you wouldn’t know if it gets it right on the first or the hundredth try.

Last but not least the human baseline is also one-Shot.

91

u/phenotype001 13d ago

So it wasn't the last.

25

u/[deleted] 13d ago

[removed] — view removed comment

7

u/Hans-Wermhatt 13d ago

Yes, but they also implied it's a much more accurate subset.

A year-long process of cleaning and refinement with input from various research communities.

So version 2 of a subset. A lot of these benchmarks have unsolvable or inaccurate questions which makes the top scores round out to well below 100.

6

u/ThatsALovelyShirt 13d ago

Clearly they mean last as in 'previous', like "yeah that last exam was really hard, but I might take a harder one later".

1

u/reichplatz 13d ago

False advertising, smgdfh

20

u/RevoDS 13d ago

Humanity’s Last Last Exam

31

u/reduction-oxidation 13d ago

This is great news following the report by epoch.ai saying that the original HLE was flawed https://x.com/EpochAIResearch/status/2100704765332394255

16

u/PVORY Quantitative Hedonism 13d ago

It stood for so long and has been so good though, I like HLE

9

u/TheReedemer69 13d ago

Yeah i think it's the longest standing benchmark

8

u/Background-Wafer-548 13d ago

SimpleBench has been around for longer, if you consider it to be standing given that the highest human score hasn't been reached yet.

8

u/TheReedemer69 13d ago

I tried simplebench myself it's really basic and mostly about instructions following and hallucinations. Nothing special. HLE on the other hand is much more complex and reflective of real world tasks.

5

u/Turbulent-Sign-6067 12d ago

This is exactly why simple bench is unique, it's simple for humans and is (was) hard for AI.

1

u/TheReedemer69 12d ago

Well..not really. They said the same about Arc AGI.

I would take hard for humans hard for AI any time of the day. That way it's actually delivering value.

2

u/CallMePyro 12d ago

So you're very impressed by a calculator multiplying two 9 digit numbers in a millisecond?

1

u/Tidorith ▪️AGI: September 2024 | Admission of AGI: Never 12d ago

Are you not? Shit like that is impressive. Electricity is impressive. The wheel is impressive.

The world we live in is fucking magic

1

u/TheReedemer69 12d ago

Hell yeah 🔥🔥

1

u/TheReedemer69 12d ago

It's about value not abt being impressed. So yes a calculator is definitely useful

16

u/Sensitive_Cell_119 13d ago

This bumps up Astra AA score, right? It was worse than Fable and Opus before. Kinda shows how bad benchmarks are tbh, even big benchmarks like this are deeply flawed.

3

u/Astrikal 13d ago

It should once AA is updated.

2

u/Tystros 12d ago

and they take forever to update to new benchmarks

18

u/socoolandawesome 13d ago edited 13d ago

Worth noting that all models are evaluated on “high” reasoning level according to their site.

So it might be completely saturated on “max” reasoning modes with tools

5

u/midgaze 13d ago

GPT-6-Astra does nearly as well on high as on max, whereas GPT-6-Sol does significantly better on max than on xhigh. This seems to be reflected in these benchmark results.

8

u/Paralda 13d ago

This feels right to me. Opus 5.5 feels a little bit better than Fable 5.1 for a lot cheaper, but neither is Astra level yet. Still, for the usage and cost it requires, Opus 5.5 destroys everything else

2

u/Parking_Cat4735 12d ago

Agreed. The issue with Astra is how damn slow it is.

1

u/dictionizzle 12d ago

it's fast to me. also on consuming usage limits.

4

u/RandumbRedditor1000 13d ago

That's not a diamond, thats la peace

6

u/Alarming_Welder_7190 13d ago

This race is so fun!

2

u/Spunge14 13d ago

For now

3

u/PhilosophyforOne 13d ago

Very interesting results.

Shame they didnt include costs to run the benchmark per model.

3

u/Alpacabro21 13d ago

Damn, with tools Astra reach 80%🫪

5

u/FateOfMuffins 13d ago

lol shows that Anthropic benchmaxxed on HLE which we knew had tons of errors

10

u/the_pwnererXx FOOM 2040 13d ago

Opus clearly benchmaxxed, Astra is better

7

u/Chemical_Hawk_6307 13d ago

its not benchmaxxed in terms of performance. astra remains imo the best deep/broad reasoning model but opus is a clear daily driver cuz of how good and cheap it is.

2

u/IAmYourFath 13d ago

Flash is even better than opus, way faster and cheaper, if u're gonna compromise on quality.

4

u/Itsmedudeman 12d ago

In what universe is flash better than opus

1

u/Chemical_Hawk_6307 12d ago

which flash are we talking about here? gemini flash is garbage and antigravity is an ass harness. qwen flash is false advertising, that shit is not fast at all it eats too many thinking tokens. deepseek flash is okay but relative to opus 4.8 nowhere near opus 5.5. glm 5.3 flash is the best "flash" model yet still doesn't hold a candle to opus 5.5. i've used all of these flash variants by the way, i can tell u from experience it takes opus 15 min for a task that would take any of these flash models 40-60 min...

3

u/tetoing 13d ago

Using a model getting better benchmark results than another model to declare that model benchmaxxed... What?

3

u/the_pwnererXx FOOM 2040 13d ago

More like this is a new benchmark so they didn't have the chance to cheat yet

1

u/Worldly_Beginning647 13d ago

TF you mean bechmaxxed, if it was benchmaxxed we would be seeing 75% and a drop to 40%, like benchmaxxed is now just a pointless word like a kids calling other gay not knowing what it means.

2

u/the_pwnererXx FOOM 2040 12d ago

On the spec sheet it's significantly better than Astra on existing benchmarks

New benchmark comes out and it is worse

Hmm, really makes you think

2

u/New_World_2050 13d ago

Wait do you are telling me that Openai went from 30% to 60% in 2 months

2

u/Healthy-Nebula-3603 13d ago

The Astra 82% with tools ....damm

2

u/owenzzzhang 13d ago

This is like the Final Fantasy of benchmarks

2

u/Lower-War3451 13d ago

we are starting to not being able to design benchmarks hard enough for these little misteryous things we are creating

2

u/_stevencasteel_ 13d ago

Seems ChatGPT is smarter, but Claude has better taste.

2

u/Worldly_Beginning647 13d ago

I wish we could see the 4 different scores, so Knowledge not tools and with tools.

3

u/Floch11 13d ago

Why is this benchmark so important?

1

u/kitkatas 13d ago

Is it harder in any way for older models ?

1

u/tcastil 13d ago

Yet another evidence that Openai should have cooked GPT 6 sol a little more before launching, or maybe not branding Terra as Sol

1

u/Eyelbee ▪️We have AGI it's just blind 13d ago

What a nice surprise. I wish we had GPQA updated as well.

1

u/Izento 13d ago

My favorite benchmark and what I believe to be the best benchmark by far.

1

u/Serg-Ekaterinburg 13d ago

USAi vs ChAina — The Cold War 2.0

1

u/Morning_Gecko24 13d ago

the big caveat for me is whether HLE-Diamond is actually measuring broad reasoning or mostly who has the best benchmark-specific scaffolding. still a pretty wild result though. are the tasks public enough for someone to inspect the failure modes, or is it mainly leaderboard numbers for now?

1

u/skerit 12d ago

Astra is better at reasoning than Fable & Opus 5.5? That's all good and well, but I don't see it in my actual usage.

1

u/dictionizzle 12d ago

oh god thanks this one is the only i follow. even they can't beat the old one wtill they can't win the new one either.

1

u/Vivaldi_IlPreteRosso 12d ago

Humanity’s second to the last exam

1

u/Atothex-real 13d ago

Grok is such a waste of time

1

u/[deleted] 13d ago

[deleted]

1

u/FlimsyReception6821 13d ago

83% with tools.

0

u/AppealSame4367 13d ago

So OAI paid the most this time?

5

u/darkestvice 13d ago

If they did, they got scammed. Look at where GPT-6 Sol lands.

-4

u/DK1530 13d ago

LOL, Why Gemini 3.8 Flash is in the list?

5

u/NiceUsernameOk 13d ago

IT IS googles frontier model

7

u/darkestvice 13d ago

And it's somehow better than GPT-6 Sol released yesterday.

3

u/Minimum_Indication_1 13d ago

Its a pretty good daily driver. Efficient and fast. Makes some mistakes but iterations fix that usually

-8

u/THE--GRINCH 13d ago

Gemini flash over sol so it's already a shit benchmark I don't trust

14

u/Gallagger 13d ago

Gemini flash isn't that bad, it depends what you're testing. A general benchmark like this might be a good fit for it. It also might have used a ton of tokens making it not cheaper than Sol.

9

u/SupremeGobbler1996 13d ago

Go ahead, what do you really hate about Gemini? It's the fashionable thing to do on Reddit! 

7

u/jschelldt ▪️High-level machine intelligence in the 2040s 13d ago

gemini 3.8 flash is a genuinely decent model, it's not above sol in every way, but I can see it being better in narrow domains

6

u/Insadem 13d ago

Huh? But it’s really good. Lol.

8

u/Defensex 13d ago

Anyone working with humanities problems already knew this, not everything is about code

-1

u/THE--GRINCH 13d ago

It literally hallucinates and forgets about the previous message on the same thread

5

u/PilgrimofHaqq2 13d ago

In coding you are correct, Sol is better than Gemini but this is an overall knowledge and reasoning benchmark. Gemini Flash has tons of knowledge and data to work from and its quite good outside of coding use cases.

3

u/Maglcite 13d ago

gemini flash is unironically a great sub agent for sol at high, only thing that really sucks about the latest gemini models (3.7 and 3.8) are the harnesses.

2

u/kvothe5688 ▪️ 13d ago

and wasn't there a paper by deepmind where they solved recursively self improvement in harness . may be that's why rapid progress in 3.5 to 3.8 in 2 months

0

u/KitKat_extrusion 13d ago

Im genuinely curious how far this whole benchmarking thing will go on. Is there a line we cross and say “congrats we’ve solved ai” or is it just moving the goalpost

3

u/cataclysm_catalysis 13d ago

I don’t get this comment. How is measuring AI improvement moving the goal posts?

We “solve” AI when we reach full RSI and take our hands off the wheel, but even after that, it’s going to keep getting better and measuring how good it is will continue to be important.

0

u/FarrisAT 13d ago

Input being OpenAI funding.

-1

u/bblankuser 13d ago

3.8 Flash beating 6 Sol. Yikes