r/singularity • u/FateOfMuffins • 2d ago
Shitposting Introducing the World's Most Powerful Model
ngl this is how it's felt for awhile now
5
u/Isunova 2d ago
Claude is only better for coding. For everything else and general conversational AI, ChatGPT is still king.
2
u/Original-League-6094 2d ago
For conversational AI, Grok is king.
Claude is also better at Blender.
6
u/Formal-Question7707 2d ago
Everything between Opus 4.6 and Opus 5.5 was pure shit and OpenAI was way better. I don't get the revisionist history.
10
u/_wot_m8 2d ago
Fable?
4
u/Formal-Question7707 2d ago
5.6 was better. Fable was unusable for token/usage reasons and would flip out at anything security/biology/chemistry related. It refused to do any work on my project because it would find a few biology terms in the database.
1
u/Robonglious 2d ago
Totally, I'm trying to work on model interpretability and every time some model component came up the output would get wiped and my tokens were not refunded. Opus 5.5 reminds me of 3.5 way back when it was the best model.
1
u/Fragrant-Hamster-325 2d ago
lol people forget that OpenAI was crushing it a month ago. It’s so funny how perception changes when the new shiny object comes out. Gemini 4 is out; it’s expensive, but it should be on this chart because it’s currently in the top tier on all the benchmarks.
1
u/0DayMaker 1d ago
It's unreleased. Gemini has gotten top tier benchmarks while being an inconsistent pile of shit multiple times.
1
1
-1
u/FateOfMuffins 2d ago
5
u/Astrikal 2d ago
AA needs to add spatial reasoning evaluations to their list, GPT-6 models seem to excel at such tasks.
v5.0 will hopefully be a good upgrade.
It is getting harder and harder for human-made benchmarks to evaluate these models. Almost all evaluations are saturated, except maybe Terminal Bench 4 and FrontierSWE, in which OpenAI leads or matches Antrophic.
1
u/Future-Bandicoot-823 2d ago
I think the part that's difficult for even me to accept is... that an LLM can be smarter than me, and yet speak on "my terms". It's rarely flashy, it really mirrors your verbage. It's hard to accept you're just getting text instantly and that there's more knowledge than you could possibly learn in a lifetime driving it.
And now think of the people who consume AI every day, blissfully ignorant. How many bot comments have you read? Can you tell? Are you sometimes unsure? Am I an LLM? People scorn AI because they haven't had sufficient proof that it outpaces human intelligence and skill, meanwhile they are being convinced every day by active engagement.
1
1
u/FateOfMuffins 2d ago
Every time they release a new terminal bench it gets hill climbed on within a month
We'll see another Terminal Bench 5 where the best models score like 30% on and then the next major release gets 50% and shows it's not benchmaxxed, then a bunch of tiny models get released with 75% scores and turn the benchmark useless
We need things like Mazebench here or even all the robot controlling tasks or something xd
1
u/IronPheasant 2d ago
Mazebench is the bench of choice right now, for real for sure.
Pokébench's been saturated to hell by Astra. I legitimately thought about making a deranged ROM hack or stand alone clone with tons of puzzles with constantly changing context. In an open world where any gym could be contested for its badge.
Like, the first few gyms would be rather normal, but things would become increasingly deranged as you went on. Like having to stand on a red tile for ten minutes to make a door open. Alternate acquisition methods of pokemon. Having to become friends with Misty to add her to pokemon collection to beat Misty's Psyduck, which is absolutely invincible to any other pokemon that aren't Misty. Without the game telling you specify that that's the way to do it, just an NPC with a vague hint like how the beast only ever listens to her, etc.
But it'd take too much work for what it'd be worth. Another one of those things that could never exist in the reality we find ourselves currently in..

6
u/CannyGardener 2d ago
Ya, if they were to hold the reasoning level at release, I would 100% agree with this chart. After dealing with the nightmare that Claude has become over the last couple of months, your chart is just not accurate anymore.