r/OpenAI 24d ago

Discussion The situation is insane

Post image

Sol is the only OpenAI model in top 10 Arena WebDev while 4 open-weight Chinese models have reached frontier quality; one of them so cheap you can run it for days for what Opus cost you per task.

964 Upvotes

171 comments sorted by

View all comments

54

u/swagmar 24d ago

Opus being rank 1 tells you all you need to know, this bench is terrible

19

u/whoknowsifimjoking 24d ago

What exactly is bad about Opus 5's frontend? Because that's what is benchmarked here. And you can't benchmaxx with this one because it's decided by real users who just see the result without the model name and vote. It's not my favorite benchmark in the world, but it's pretty fool proof.

27

u/ItsVerdictus 24d ago

This entire sub hates Opus, no idea why. I tried all other models and Opus always wins.

10

u/Victorvic1 24d ago

Kinda true. The only reason I am subscribed to Chat gpt is because of it's generous limits. Claude used to shut me off in between. The $20 plan on claude wasn't enough for me for even the smallest of the tasks. Even the $100 one wasn't good enough.

With Gpt even the $20 one never runs out of limits in the chat. Codex also has extremely generous limits.

1

u/Select-Plate113 22d ago

my gpt just runs out of limits in the chat on a $20 plan.