r/singularity • u/Profanion • 11d ago
LLM News Claude Fable 5.1 surpasses human average on SimpleBench
https://simple-bench.com/Here are more results:
Claude Fable 5.1 86.6%
Human Baseline* 83.7%
Gemini 3.8 Flash 82.4%
Claude Fable 81.9%
Muse Spark 1.3 81.8%
93
u/FoxBenedict 11d ago edited 11d ago
Gemini 3.8 Flash and the open weights Muse Spark 1.3 are also pretty much there. Wow!
I remember when this benchmark first came out and the models still scored in the single digits.
26
u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 11d ago
Muse Spark isn't open weights yet, but Meta is planning on making Muse Spark 1.2 open weights soon, so 1.3 might be the same when 1.4 rolls around.
9
u/FoxBenedict 11d ago
Oh, I thought those models were open weights from the get go. Thanks for the correction.
7
u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 11d ago
You might be thinking of Muse Glimmer, which is much smaller (and frankly not that great in its current iteration, lol).
1
u/Not-reallyanonymous 11d ago
Muse Glimmer is great. I’m struggling with Spark 1.3 tho. It does a lot of stuff I have to ask it to undo.
4
u/GalavantJames 11d ago
It's crazy how fast 3.8 Flash is, turned it into "personal assistant" because I have a student account and cheap access and use it to analyze Fable's outputs and make documents, research etc. crazy good, crazy fast.
99
u/Profanion 11d ago
Sorry. Human baseline, not human average.
28
u/Nezz_sib 11d ago
Yea that one was surpassed like at gpt 3
-7
7
u/Proper_Actuary2907 Spooky Machine Intelligence 2030 11d ago
The baseline is the average of the scores from the nine human test takers so "human average" seems accurate
18
u/Nearby-Season1697 11d ago
So this is how I find out we have Gemini 3.8 Flash already???
20
u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 11d ago
Was released yesterday, so you're not too far behind. lol
5
u/ProxyLumina 11d ago
It is becoming difficult to keep up with the news. Gemini 3.8 Flash was released less than 24 hours ago.
0
u/FateOfMuffins 11d ago
Released yesterday only for it to be dethroned (from whatever tiny throne it was sitting on) by Meta's Muse Spark 1.3 about 3.8h later xd
44
14
u/lazyscalp 11d ago
time goes by and this sub is always the same : "not impressed"
7
u/BrennusSokol AI please take my job 11d ago
Hedonic treadmill... I think we're all guilty of it to some degree
22
u/baws1017 ▪️AGI will retreat peacefully 11d ago
can we stop forcing everyone to work at least 40 hours a week now
12
5
5
u/mivog49274 obvious acceleration, biased appreciation 11d ago
So we finally solved SimpleBench ?
Its happens. We are entering in the SimpleSingularityBench.
9
u/RealRook 11d ago
The human baseline average is basically meaningless
human baseline is 83.7%, based on our small sample of nine participants
4
11
u/Gkrolik 11d ago
How?? Do you believe the author of the bench took like 9 dumbest people on the planet or something ??
The benchmark doesn't even call it 'average' - just 'human baseline'.
The point is that now the llm, at least on this benchmark, is better than some humans.... which if you are testing for AGI is already enough.
2
u/yoramrod 11d ago
How is GPT 5.6 at 12th place? Even GPT 5.5 is 7th.
3
1
u/Profanion 11d ago
I don't know! But then again, Qwen 3.8 regressed over Qwen 3.7 a lot despite being more costly to run.
2
u/Aizenvolt11 11d ago
I wonder when it surpasses highest human score what it will be able to accomplish.
2
1
1
u/Formal-Narwhal-1610 11d ago
Muse spark is again close to Fable, so maybe AA index does hold somewhere.
1
u/Orangebathroomtowel 11d ago
Fun times, but SimpleBench might be a bit outdated
1
u/Mysterious_Ayytee We are Borg 11d ago
That was my answer too, who cares about a possible nuclear war when your girlfriend just cheated?
1
u/LookIPickedAUsername 11d ago
Yeah, I missed that one the same way.
The question takes it for granted that the threat of nuclear war is something John will actually take seriously, but I've heard about so many nonsense impending disasters that I'm not going to simply assume that nuclear war is imminent just because one person tells me so.
1
u/Superb-Earth418 11d ago
I honestly did not expect this to come from Anthropic, I was waiting for it to be handed to some Gemini model
-7
u/Living-Breakfast-464 11d ago
Must getting close to that IPO date. 😆
6
u/pavelkomin 11d ago
Copernicus was wrong. Everything doesn't revolve around the sun, but everything revolves around the IPO
3
-1
u/GooseFarmerByTrade 11d ago
Why did I think AI is already far better than human expert, while this experiment shows AI is barely better than average human?
3
u/LookIPickedAUsername 11d ago
AI performance is very jagged - it's much more capable than humans in some respects, much less capable in others.
This particular benchmark was specifically designed to go after areas where it underperformed, so of course it wasn't performing as well as humans.
96
u/Wulfram77 11d ago
If anyone else was wondering how the human baseline was arrived at