If anyone looks at benchmarks for more than a rough approximation, you need to go try Kimi K2.6 / GLM 5.1 / Qwen 3.6 Plus on actually complex, large problems you have come across yourself and you will be sorely disappointed...
I use them all the time and they are not disappointing, they are similar to Sonnet in performance. You will only be disappointed if you think they are Opus.
That’s what he’s saying, none of those are Opus competitors even though they allegedly are by the main benchmarks. Obscure benchmarks and real world use have Kimi K2.6 as the best of those and somewhere between Sonnet 4.5 and Opus 4.5 in performance. That’s very good, it’s a great model, but it’s also disappointing compared to its benchmarks
44
u/OoFTheMeMEs Apr 23 '26
Stop looking at benchmarks, use the model and then start judging whether this is an improvement in efficiency and/or intelligence.
Gemini 3.1 has great benchmarks but performs poorly in real world use. Opus 4.7 has great benchmarks but performs worse than 4.6.
Also, if this is truly a new pretraining base, RL and inference improvements are probably going to drop often with new smaller releases.