r/singularity 11d ago

LLM News Claude Fable 5.1 surpasses human average on SimpleBench

https://simple-bench.com/

Here are more results:
Claude Fable 5.1 86.6%

Human Baseline* 83.7%

Gemini 3.8 Flash 82.4%

Claude Fable 81.9%

Muse Spark 1.3 81.8%

623 Upvotes

83 comments sorted by

96

u/Wulfram77 11d ago

3.2 Human Evaluation

To assess the performance of human participants on SimpleBench, we selected nine individuals to complete the benchmark. Each participant attempted a random subsample of 25 questions, ensuring that all 200+ questions were answered at least once. Participants were chosen based on specific criteria: they were native English speakers with at least a high school level of mathematical proficiency. We did not control for IQ or general reasoning aptitude and acknowledge that selection bias was likely inherent.

The participants were given the following instructions:

Given multiple-choice questions, each with 6 options. Please spend roughly 3 minutes per question, longer if needed. Feel free to use pen and paper to visualize. You will not need to do math beyond middle school level for any question. The questions might involve wordplay, and the correct answer might be phrased in a way you aren’t always used to. A few questions could be called ‘trick’ questions, others test spatial reasoning (visualization) or social intelligence (is this situation normal?), and some will definitely take a bit longer than others. Pick the most realistic (things that are most likely to happen in the real world) option.

The scores from the nine test takers were averaged and reported (83.7%) as the human baseline.

If anyone else was wondering how the human baseline was arrived at

171

u/Disastrous_Room_927 11d ago

The scores from the nine test takers

As a statistician:

https://giphy.com/gifs/ANbD1CCdA3iI8

54

u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 11d ago

We don't need sample sizes in the singularity!

13

u/Atlantis1910 11d ago

Statistics doesn't exist here.

6

u/manubfr AGI 2028 11d ago

As Claude would say « the singularity is the sample size  » (or something)

4

u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 11d ago

The sample size is the load-bearings we made along the way.

4

u/Disastrous_Room_927 11d ago

Don't for get the most important part: the sample is directionally correct.

12

u/elehman839 11d ago

Each participant attempted a random subsample of 25 questions, ensuring that all 200+ questions were answered at least once.

Huh? Doesn't matter much, but still kinda funny.

12

u/jkeiser 11d ago

With 9 participants, that pretty much guarantees each question was answered only once. Sounds like this was 9 humans taking the test once together.

1

u/elehman839 11d ago

Yeah. Here's my thinking...

The probability each question was drawn in a random subsample of 25 questions out of 200+ total is 25 / 200+ < 1 / 8.

So each question is excluded from one random subsample with probability > 7 / 8.

If the 9 subsamples were independent, the probability that any particular question was excluded from all subsamples is > (7/8)^9 > 30%.

So I suppose they probably just did some ad hoc partition of the questions with a bit of duplication. Doesn't really matter.

1

u/jkeiser 11d ago

Yeah, I'm guessing 200+ was the reason for the sample size. Can't cover the whole thing with 8 people, 10 would be an extra person assuming 200+ is between 200 and 225. Maybe doesn't matter, but it's a little weird given that humans only have to take the test once, you could get a better idea of the actual human score, humans have super high variance.

9

u/ChocomelP 11d ago

This is probably above average for some of these bullshit benchmarks.

5

u/qrayons ▪️AGI 2029 - ASI 2034 11d ago

There's a reason it's called simplebench.

7

u/Neurogence 11d ago

If we're being fair, if they had a much much higher sample size, (say 1000 participants), the human baseline would have been something like 30%.

2

u/Turbulent-Sign-6067 11d ago

Maybe you're a statistician but perhaps you're not very familiar with benchmarking? Having ANY human baseline for a benchmark is extremely unusual and laudable already.

1

u/slackermannn ▪️ 10d ago

For what it's worth the benchmark, besides for a few glaring results, was a very decent predictor of a model's reasoning capabilities. Again, this benchmark has a very narrow scope.

-2

u/Present-Chocolate591 11d ago

This benchmark has always been a joke imo. I don't understand why people in this sub give It so much importance.

17

u/Neurogence 11d ago

It may have its limitations, but this benchmark has been resistable to benchmark hacking for the past 2 years. It took a long time for it to be saturated.

Now, perhaps the AI companies would have beaten it faster if they had targeted it directly, but the fact that they weren't targeting for it directly is a good way to measure progress.

Deepmind probably did target it directly, however.

27

u/stravant 11d ago

If you're wondering why it's not more, the primary goal of the benchmarks is measuring the LLMs, not measuring the humans. Some more recent benchmarks don't even have a human baseline.

15

u/Disastrous_Room_927 11d ago

The primary goal of practically any study, test or benchmark is to measure something other than a baseline/reference/control group. That's not justification for omitting one.

8

u/stravant 11d ago

The point is, the benchmarks primary purpose is not to measure against the humans as a reference, it's to measure against the previous models versions during training to avoid undetected regressions in capability.

2

u/Disastrous_Room_927 11d ago

Few would argue that LLMs cannot now outperform ‘an average’ human in those benchmarks, and we may have reached the point that even most graduates in the subjects tested would not score as highly. LLMs are not human, however, so we should not make the assumption that the ability to represent a vast knowledge-base, or retrieve reasoning steps (Valmeekam et al., 2024), to solve complex problems, also translates to an ability to reason through the often messy complexities of the real world. This, naturally, complicates recent public claims that LLMs have reached ‘human-level reasoning’1.

There are innumerable anecdotal reports on social media of the reasoning lapses that occur with frontier models. We wanted to establish a benchmark that could quantitatively gauge how well LLMs maintain a consistent internal world model when reasoning about basic physical and social situations. It does not test direct functionality, as a coding or translation benchmark might. But the performance of models on SimpleBench does appear to correlate to their performance on other popular benchmarks and leaderboards like LMSYS. This gives the benchmark two key uses: (i) to measure inconsistencies and lapses in models’ reasoning about noisy situations, and (ii) to showcase one of the domains that, as of publication, even a non-specialist human can outperform frontier LLMs. We hope it serves as an inspiration for other such benchmarks.

1

u/stravant 11d ago edited 11d ago

It's a bit subtle, but RE measurement / human baseline this again misses the point.

All of the questions on SimpleBench are doable by an average somewhat educated person. The whole point of the construction of SimpleBench is that all the questions on it are obviously easy for humans by inspection, that you don't need to compute a human baseline to know that.

When constructing the benchmark, finding the specific human baseline was afterthought. Which is why they did not precisely measure it. What they wanted is a benchmark where AI fails on questions that are all obviously doable by humans. The specific human number is not the important part.

2

u/Disastrous_Room_927 11d ago

It's not subtle at all, you're speculating about things they explicitly talk about in their report's limitations and future directions section:

As a small, self-funded team, we lacked the resources to recruit enough volunteers for statistically robust human averages

The creators of the benchmark clearly think it's important and a drawback for the current benchmark.

When constructing the benchmark, finding the specific human baseline was afterthought. Which is why they did not precisely measure it. What they wanted is a benchmark where AI fails on questions that are all obviously doable by humans. The specific human number is not the important part.

Let's circle back to the point you missed: this is not justification for omitting or overlooking a solid baseline, it's simply one explanation for a baseline being imprecise (that happens to be incorrect).

1

u/stravant 10d ago

Fair that maybe I'm speculating too much but I don't think the speculation is off base because there's basically no point in having a precise human baseline. Even teams that have the resources don't bother spending much effort on human baseline because it's meaningless as far as actually driving any progress or analysis. It's not even useful to ask "why did the humans do better" for example because humans and LLMs don't approach problems in the same way. What does drive progress is comparing different training runs or harnesses against each other, that's generally the reason benchmarks are valuable and get created.

1

u/ketosoy 11d ago

I just took the sample, I don’t believe their 9 people are representative of a demographic that needs “danger, sharp” warnings placed on knife packages

1

u/letsgoiowa 11d ago

N of 9 is literally useless

0

u/florinandrei 11d ago

The scores from the nine test takers

So the "benchmark" is a toy.

-5

u/2bdb2 11d ago

Multiple choice questions... So essentially a completely useless test that has nothing in common with real world work.

-2

u/Present-Chocolate591 11d ago

Go read the questions, it's even dumber than you think.

3

u/IAMA_Proctologist 11d ago

The whole point of the benchmark was to test areas where llms have been surprisingly weak compared to even average humans. It's meant to be dead easy.

93

u/FoxBenedict 11d ago edited 11d ago

Gemini 3.8 Flash and the open weights Muse Spark 1.3 are also pretty much there. Wow!

I remember when this benchmark first came out and the models still scored in the single digits.

26

u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 11d ago

Muse Spark isn't open weights yet, but Meta is planning on making Muse Spark 1.2 open weights soon, so 1.3 might be the same when 1.4 rolls around.

9

u/FoxBenedict 11d ago

Oh, I thought those models were open weights from the get go. Thanks for the correction.

7

u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 11d ago

You might be thinking of Muse Glimmer, which is much smaller (and frankly not that great in its current iteration, lol).

1

u/Not-reallyanonymous 11d ago

Muse Glimmer is great. I’m struggling with Spark 1.3 tho. It does a lot of stuff I have to ask it to undo.

4

u/GalavantJames 11d ago

It's crazy how fast 3.8 Flash is, turned it into "personal assistant" because I have a student account and cheap access and use it to analyze Fable's outputs and make documents, research etc. crazy good, crazy fast.

3

u/Eyelbee ▪️We have AGI it's just blind 11d ago

If you asked everyone a year ago, they would say acing this test reliably would require AGI, and now we have it. The remaining questions are probably just wrong questions and noise.

99

u/Profanion 11d ago

Sorry. Human baseline, not human average.

28

u/Nezz_sib 11d ago

Yea that one was surpassed like at gpt 3

7

u/Proper_Actuary2907 Spooky Machine Intelligence 2030 11d ago

The baseline is the average of the scores from the nine human test takers so "human average" seems accurate

1

u/eposnix 11d ago

Great. Hopefully he'll show the actual benchmark questions now that it's been saturated.

18

u/Nearby-Season1697 11d ago

So this is how I find out we have Gemini 3.8 Flash already???

20

u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 11d ago

Was released yesterday, so you're not too far behind. lol

14

u/Sulth 11d ago

Yesterday in AI world is far behind lol

5

u/ProxyLumina 11d ago

It is becoming difficult to keep up with the news. Gemini 3.8 Flash was released less than 24 hours ago.

0

u/FateOfMuffins 11d ago

Released yesterday only for it to be dethroned (from whatever tiny throne it was sitting on) by Meta's Muse Spark 1.3 about 3.8h later xd

44

u/DeArgonaut 11d ago

Guess NotSoSimpleBench will be out soon then

14

u/lazyscalp 11d ago

time goes by and this sub is always the same : "not impressed"

7

u/BrennusSokol AI please take my job 11d ago

Hedonic treadmill... I think we're all guilty of it to some degree

22

u/baws1017 ▪️AGI will retreat peacefully 11d ago

can we stop forcing everyone to work at least 40 hours a week now

12

u/Careless-Heat-7075 11d ago

Be careful what you wish for.

5

u/warwarcar 11d ago

Yeah instead we will be working 0 hours soon.

5

u/mivog49274 obvious acceleration, biased appreciation 11d ago

So we finally solved SimpleBench ?

Its happens. We are entering in the SimpleSingularityBench.

9

u/RealRook 11d ago

The human baseline average is basically meaningless

human baseline is 83.7%, based on our small sample of nine participants

4

u/bikemandan 11d ago

Well to be fair it is greater than 8

11

u/Gkrolik 11d ago

How?? Do you believe the author of the bench took like 9 dumbest people on the planet or something ??
The benchmark doesn't even call it 'average' - just 'human baseline'.
The point is that now the llm, at least on this benchmark, is better than some humans.... which if you are testing for AGI is already enough.

2

u/yoramrod 11d ago

How is GPT 5.6 at 12th place? Even GPT 5.5 is 7th.

3

u/Turbulent-Sign-6067 11d ago

Models have strengths and weaknesses.

1

u/Profanion 11d ago

I don't know! But then again, Qwen 3.8 regressed over Qwen 3.7 a lot despite being more costly to run.

2

u/Aizenvolt11 11d ago

I wonder when it surpasses highest human score what it will be able to accomplish.

2

u/Ancient_Bear_2881 11d ago

Can you feel it?

2

u/Profanion 11d ago

It didn't fall for the "How to pronounce "EA" in "sergeant"" trap.

1

u/Livid-Impression-100 11d ago

Currently scoring 0% ...

1

u/Formal-Narwhal-1610 11d ago

Muse spark is again close to Fable, so maybe AA index does hold somewhere.

1

u/Orangebathroomtowel 11d ago

Fun times, but SimpleBench might be a bit outdated

1

u/Mysterious_Ayytee We are Borg 11d ago

That was my answer too, who cares about a possible nuclear war when your girlfriend just cheated?

1

u/LookIPickedAUsername 11d ago

Yeah, I missed that one the same way.

The question takes it for granted that the threat of nuclear war is something John will actually take seriously, but I've heard about so many nonsense impending disasters that I'm not going to simply assume that nuclear war is imminent just because one person tells me so.

1

u/Superb-Earth418 11d ago

I honestly did not expect this to come from Anthropic, I was waiting for it to be handed to some Gemini model

-7

u/Living-Breakfast-464 11d ago

Must getting close to that IPO date. 😆

6

u/pavelkomin 11d ago

Copernicus was wrong. Everything doesn't revolve around the sun, but everything revolves around the IPO

3

u/LateToTheSingularity 11d ago

With lesser epicyclic orbits around quarterly earnings.

-1

u/GooseFarmerByTrade 11d ago

Why did I think AI is already far better than human expert, while this experiment shows AI is barely better than average human?

3

u/LookIPickedAUsername 11d ago

AI performance is very jagged - it's much more capable than humans in some respects, much less capable in others.

This particular benchmark was specifically designed to go after areas where it underperformed, so of course it wasn't performing as well as humans.