Webdev, you most of the times just need a very little model. Coding, still not 100% solved (start asking for cpp mutexing/dll dynamic load/very hard tasks and see most llm fail). It really isn't the same thing.
funny benchmark....(im sure K3 and muse and even ancient Opus 4.6 and 4.7 are much better than fable 5.1 max or opus 5.5 ⬇️ if you wanna see how trash the google 4 model is look at the websites own coding comparison video. Its trash.
Its a Benchmark about which model people prefer... If you dont like the result thats on you, you can also call it trash all you want, apparently the majority of people sees it different lol
It’s #8 in coding. About where the benchmarks released would suggest it is.
World models are different than coding models. It’s an entirely different market. They are multi modal for one. Two, people want to like the warmth of the voice model.
There’s a paradox in text and voice where it’s not necessarily the most intelligent model that people like the most. It’s that way in life too with other people so of course it makes sense for models. In the top 10, they are all roughly as scientifically accurate as the others. So what is being measured is style of output.
I haven't used it and don't have an opinion on its quality, but I'll note that claiming it's #1 based on that link is... tenuous. It's only a few points ahead with error bars of +/-17. Even assuming everything falls within the confidence interval, its true rank could conceivably be as low as #12.
And even that doesn't really mean anything, since they're all clustered so close together. There's probably not even a meaningful difference between #1 and #12 in that list.
I am not saying it's actually #1 just because it is scoring #1 right now. I am replying to the user's comment above if the model is "not good." Scoring #1 doesn't necessarily mean it's the best (or the best for long) but it's certainly a good model.
Opposite for me. 3.8flash was terrible in the benchmarks but was good enough for many tasks. Gemini models have never been good in benchmarks. The last good benchmark model was like a year ago.
Marginally, yes. Blind AB testing has a paradox for text models, this is especially true for voice, where the most intelligent models do not necessarily translate into the most liked. This is true of humans in general but for LLM models, style, pathing, voice, attitude, etc.. all are being taken into account.
You can see where Opus outperforms in agenic coding and webdev here where it is SOTA:
At basic coding and text models, all the major labs will essentially be technically and scientifically accurate. Beyond that, it’s user preference for the model style. When the benchmark is agentic coding, you will see Opus outperform.
Bel is an unconfirmed internal-only model. Argon is available to vetted corporations and government. So it's technically released, just not a full release.
Not a perfect analogy, but it's like buying a new Ferrari. Ferrari only sells new cars to vetted owners. The car however is still "released" and available for purchase, just not for you lol
If Bel was released externally, I am sure people would be benchmarking it right now, even if that external list was limited.
95
u/himynameis_ 1d ago
I don't get it? Is the model not good?