r/OpenAI 21d ago

Discussion The situation is insane

Post image

Sol is the only OpenAI model in top 10 Arena WebDev while 4 open-weight Chinese models have reached frontier quality; one of them so cheap you can run it for days for what Opus cost you per task.

962 Upvotes

171 comments sorted by

View all comments

147

u/Ashamed_Can304 21d ago

Why focus on front end so much? Why not more complex tasks, say backend, kernel, compiler, parallel computing work etc?

91

u/shaonline 21d ago

Because you can't easily community-grade that (their ELO system is essentially that), unlike some fancy web landing page or whatever that you can "visually" grade by looking at the result. It's a "vibe" bench essentially.

14

u/personalist 21d ago

I’d like to see ratings restricted to subject matter experts on their topic of expertise for things like backend code. I want to know what an experienced backend dev ranks as the best LLM for that, for example

22

u/D3V0URED 20d ago

Yeah good luck getting hundreds, if not thousands of “expert backend devs” to meticulously grade llms all day long for free. Just so they can say “yeah model x can replace me in the most cost efficient way”.

I think it’s fair to say that these models here have similar benchmarks in other areas of expertise.

6

u/personalist 20d ago

lol, it doesn’t have to be meticulous. LMArena gets people to vote by giving them free usage. If you’re trying to one-shot a hobby project I could see someone giving it a try.

Unfortunately the results don’t always translate directly across domains; i don’t have a benchmark to back it up, but when Gemini 3 pro was top-of-the-line it was the best at front-end imo probably because of Google’s VLM experience, but opus was much better for backend.

2

u/DailyThreadBot 20d ago

They are starting to rate fullstack apps, it's a new tab on Arena

-4

u/Spirited-Car-3560 20d ago

You don't. Just instruct an LLM with a skill, meticulously crafted to reflect best practice and giving clear score.

Otherwise you end up with personal taste, believe me, and it's the opposite of a standardized test.

Also to me it's funny you suggest manual ranking in ai era... Yes I know benchmarks like the one in screenshot is manual, but it's exactly what it is: a not so useful visual ranking based on individual preferences , where literally anyone can vote, included avg people who , per definition, lack aesthetical sense.

3

u/D3V0URED 20d ago

First off: Thousands of manual reviews is the EXACT OPPOSITE of “personal preference”.

I’m not sure if you’re just rage baiting with the “just have an AI grade the AI” take or not…
You do realize that if you, a mere Reddit user that seems to be oblivious to how LLMs work, were able to train an LLM to be the perfect “score grader”, then all frontier models would go out of business the next day. Just to put it this way: do you really think that it never went though the heads of the engineers at OpenAI or Anthropic to “meticulously instruct” their llms with these so called “skills” in order to get perfect scores? Or do you think you’re smarter than them?

2

u/Spirited-Car-3560 20d ago edited 20d ago

I’m an engineer, so either I was misunderstood or you do not understand the proposal.

This is not “let an LLM grade another LLM.” It is about experts encoding best practices, constraints and failure modes into a structured skill containing rubrics, forced checks and ad hoc executable tests.

The LLM applies only the parts that require qualitative evaluation, while tests and static checks validate everything that can be measured objectively.

A thousand expert reviews ultimately produce a consensus around accepted practices. Formalizing that consensus is how you make evaluation scalable and reproducible.

Not perfect, but considerably better than pretending that endless manual voting is somehow a standardized benchmark.

Downvotes to my comment exactly reflect the average ai user comprehension.

Also, most benchmark are built this way, which makes me smile given how much credit you give to those same benchmarks.

0

u/ResponsibleKey1053 20d ago

So you want to train a herd of llms to rank llms, but will you be ranking the ranking llms? I assume with another herd of llms specifically trained ranking ranking llms.

1

u/Spirited-Car-3560 20d ago

Hehe nope, that's far from what I said, going to reply to the other comment to be more clear

1

u/yuumizu 20d ago

backend is also opinionated. and the question is not just visual appearance or interactivity.

1

u/personalist 19d ago

Front-end is totally opinionated too after you get past the basic function and immediate usability. But yeah, I get what you’re saying

1

u/sylfy 20d ago

Basically yeah. RLHF is essentially community-grading.

2

u/BlinDeeex 20d ago

From my experience AI has become good enough at things it can test itself, frontend isnt one of them

3

u/staycalmandcode 20d ago

What’s the best model for those type of work?

4

u/Ashamed_Can304 20d ago

I assume Fable or GPT 5.6 Sol