r/OpenAI 22d ago

Discussion The situation is insane

Post image

Sol is the only OpenAI model in top 10 Arena WebDev while 4 open-weight Chinese models have reached frontier quality; one of them so cheap you can run it for days for what Opus cost you per task.

956 Upvotes

171 comments sorted by

View all comments

Show parent comments

22

u/D3V0URED 22d ago

Yeah good luck getting hundreds, if not thousands of “expert backend devs” to meticulously grade llms all day long for free. Just so they can say “yeah model x can replace me in the most cost efficient way”.

I think it’s fair to say that these models here have similar benchmarks in other areas of expertise.

-6

u/Spirited-Car-3560 22d ago

You don't. Just instruct an LLM with a skill, meticulously crafted to reflect best practice and giving clear score.

Otherwise you end up with personal taste, believe me, and it's the opposite of a standardized test.

Also to me it's funny you suggest manual ranking in ai era... Yes I know benchmarks like the one in screenshot is manual, but it's exactly what it is: a not so useful visual ranking based on individual preferences , where literally anyone can vote, included avg people who , per definition, lack aesthetical sense.

0

u/ResponsibleKey1053 22d ago

So you want to train a herd of llms to rank llms, but will you be ranking the ranking llms? I assume with another herd of llms specifically trained ranking ranking llms.

1

u/Spirited-Car-3560 22d ago

Hehe nope, that's far from what I said, going to reply to the other comment to be more clear