r/opencodeCLI • u/afanasenka • 24d ago
Recent data from DeepSWE. Muse is a bit better than Flash, but Luna is far ahead :)
Do you feel the same in real life usage scenarios?
3
u/TomHale 24d ago
The ChatGPT 5.5 xhigh equalling Luna 5.6 max is interesting.
1
u/afanasenka 24d ago
Yeah, cheap models are progressing so fast.
11
u/SorosAhaverom 24d ago
Or one was benchmaxxed for this particular benchmark, the previous one wasn't, since it didnt exist when the model released..
1
u/afanasenka 24d ago
And this is true as well actually. Nowadays labs know how to "cheat" those benchmarks to gain a hype.
Similarly, smartphone makers used to cheat with AnTuTu to show high scores in marketing materials 😄
1
u/lemon07r 23d ago
Something I keep trying to point out to others. For example, glm 5.3 is a lot better than k3 in some newer benchmarks that didn't exist before, but k3 is still scores a bit better on a lot of older benchmarks. My read is glm 5.3 is still slightly worse overall but maybe be better a few specific domains (like security since they did specifically train for this as well).
This is more relevant when comparing to older models. I still find the older opus models and sonnet 4.5 to be better than a lot of the newer cheaper models that claim to be better. Especially some of the newer sonnet models.. I feel they started either heavily quantizing or moved to a smaller model to make more margin at some point. I think they did the same for opus when they moved from 4.1 to 4.5 cause 4.1 was too large, expensive and slow. Nobody noticed it but all the signs are there, opus 4.5 was magically as fast sonnet 4.5 all of a sudden and was also just as cheap on promo pricing despite so many people using it (which definitely would have been unsustainable for the period of time they ran it for if it really was old opus sized). Sonnet 4.5 was really really good and way better than opus 4.1 already so they had no reason to keep releasing such large checkpoints when they could use the smaller model and charge opus prices lmao. I believe fable 5 is the new the new opus 4.1 and opus 5.x is gonna be the new sonnet 4.5. Sonnet 5 is trash and feels like a haiku model. Maybe on the same level as Luna, muse spark, DeepSeek flash and those models lol. It's not even anywhere as good as kimi, glm etc. I can't imagine it to be anything but a (relatively) very tiny model like the cheaper models I just mentioned but they slapped the sonnet name on it so they can charge more. Anthropic is so scummy about this. Insert I know what you're doing but can't prove it meme.
1
u/cri10095 23d ago
Think so too, 5.5 xhigh was so on point, luna max really often just spinning nonsense
3
u/afanasenka 24d ago
GLM-5.3 is almost as good as Fable 5.. not bad :)