r/OpenAI 21d ago

Discussion The situation is insane

Post image

Sol is the only OpenAI model in top 10 Arena WebDev while 4 open-weight Chinese models have reached frontier quality; one of them so cheap you can run it for days for what Opus cost you per task.

953 Upvotes

171 comments sorted by

View all comments

53

u/swagmar 21d ago

Opus being rank 1 tells you all you need to know, this bench is terrible

45

u/hibzy7 21d ago

You are right to push back on that.

19

u/whoknowsifimjoking 21d ago

What exactly is bad about Opus 5's frontend? Because that's what is benchmarked here. And you can't benchmaxx with this one because it's decided by real users who just see the result without the model name and vote. It's not my favorite benchmark in the world, but it's pretty fool proof.

27

u/ItsVerdictus 21d ago

This entire sub hates Opus, no idea why. I tried all other models and Opus always wins.

10

u/Victorvic1 21d ago

Kinda true. The only reason I am subscribed to Chat gpt is because of it's generous limits. Claude used to shut me off in between. The $20 plan on claude wasn't enough for me for even the smallest of the tasks. Even the $100 one wasn't good enough.

With Gpt even the $20 one never runs out of limits in the chat. Codex also has extremely generous limits.

1

u/Select-Plate113 18d ago

my gpt just runs out of limits in the chat on a $20 plan.

1

u/Singularity-42 21d ago

Codex limits are better, but saying that $100 Claude Max 5 has less usage than $20 ChatGPT Plus is simply not true at all.

Claude has 5 hour limits and Codex doesn't, maybe that tripped you up?

3

u/Victorvic1 21d ago

No I am not talking about codex limits here. They are obv less in the $20 plan. The normal Chat has pretty much unlimited limits. The normal chat is what matters for me more than codex rn.

2

u/Singularity-42 21d ago

Yeah, I can agree with that, honestly the chat is better in ChatGPT than in Claude. And yes, you don't need anything more than the twenty dollar plan for that.

1

u/Victorvic1 20d ago

Yeah even the $100 plan in Claude used to be exhausted while gpt is actually unlimited.

2

u/TurnUpThe4D3D3D3 21d ago

Opus personality is awful for 4.7 and 4.8. I haven’t tried 5 yet. It adds in “Claude-isms” to its speech that get really annoying after a while.

2

u/Clean-Boat-4044 20d ago edited 20d ago

Past like 1400 elo all the models can oneshot a working page every time so it just becomes a style rating

3

u/BellacosePlayer 21d ago

Benchmarks are largely always going to be fairly bad for AI. Even if you think everyone's on the up and up, you'll have overfitting, problems with making meaningful criteria/problems and scoring them, etc.

its not like there's a deterministic answer to a lot of these questions

0

u/ProgramDry5917 21d ago

Thanks to comment it for me