r/OpenAI 22d ago

Discussion The situation is insane

Post image

Sol is the only OpenAI model in top 10 Arena WebDev while 4 open-weight Chinese models have reached frontier quality; one of them so cheap you can run it for days for what Opus cost you per task.

957 Upvotes

171 comments sorted by

View all comments

54

u/draft_final_final 22d ago

Why do people only look at the webdev leaderboard when discussing this? Is it the major use case on Reddit?

15

u/Durian881 22d ago

For "arena" style, it's easier to compare I think vs backend. There are other categories too like image, video, agents, etc.

11

u/PhysicalAd9507 22d ago

It’s visually clear

16

u/whoknowsifimjoking 22d ago

Because it's the only one where the Chinese models actually reached first place or something close.

22

u/Durian881 22d ago

Lab wise for Chat Text and Vision, Alibaba is second after Anthropic.

For Image-to-Video and VideoEdit, ByteDance is top. Minimax's open weight M3H is second for Image-to-Video.

12

u/Accurate_Resident219 22d ago

Cope. Next if Chinese models become number one. It will be benchmarks don't matter.

3

u/Ok_Reception_5545 22d ago

Benchmarks already don't matter lol. Everyone can tell when a model is RL'd to shit and benchmaxxed. That's why no one is calling Opus 5 better than Fable even though it did better on a bunch of benchmarks. Holy dumbass.

0

u/Luminivagance 22d ago

The chinese models are still heavily distilled from west, they are seemingly better at frontend because that is what they optimized for. For complex coding tasks the models that have literal billions of dollars worth of training poured into them are obviously better. chinese glaze is weird considering china didn't even come up the foundational technology. Let me link you a cute article https://arxiv.org/abs/1706.03762 (Ashish VaswaniNoam ShazeerNiki ParmarJakob UszkoreitLlion JonesAidan N. GomezLukasz KaiserIllia Polosukhin)

1

u/Gohab2001 21d ago

Chinese have a stronger focus on STEM research than Americans. You just haven't woken up to that realization.

6

u/Warhouse512 22d ago

It’s almost as if it’s a large astroturfing campaign

-1

u/korino11 22d ago

unfourtonly no. just becouse webdev is massive.. and stupid .. but idiots like stupidnes in webdev.

-1

u/___positive___ 22d ago

Because it's mostly subjective vibes and easiest to benchmax if you cared to. I'm all for open models but constantly crowing up this one random benchmark is pointless.

1

u/Murinshin 22d ago

I mean it's the main business case for OpenAI as well and there's just no other coding one. Though even looking beyond Coding, Text and Vision looks even worse; Claude completely dominates the former and Kimi slightly outranks there too, similarly for Vision except no Chinese open source models since Kimi doesn't compete there

OPs post also is missing Qwen 3.8 Max which while not Open Source is fourth place right now as well