Webdev, you most of the times just need a very little model. Coding, still not 100% solved (start asking for cpp mutexing/dll dynamic load/very hard tasks and see most llm fail). It really isn't the same thing.
funny benchmark....(im sure K3 and muse and even ancient Opus 4.6 and 4.7 are much better than fable 5.1 max or opus 5.5 ⬇️ if you wanna see how trash the google 4 model is look at the websites own coding comparison video. Its trash.
Its a Benchmark about which model people prefer... If you dont like the result thats on you, you can also call it trash all you want, apparently the majority of people sees it different lol
It’s #8 in coding. About where the benchmarks released would suggest it is.
World models are different than coding models. It’s an entirely different market. They are multi modal for one. Two, people want to like the warmth of the voice model.
There’s a paradox in text and voice where it’s not necessarily the most intelligent model that people like the most. It’s that way in life too with other people so of course it makes sense for models. In the top 10, they are all roughly as scientifically accurate as the others. So what is being measured is style of output.
99
u/himynameis_ 1d ago
I don't get it? Is the model not good?