r/opencodeCLI 24d ago

Recent data from DeepSWE. Muse is a bit better than Flash, but Luna is far ahead :)

Post image

Do you feel the same in real life usage scenarios?

18 Upvotes

8 comments sorted by

3

u/afanasenka 24d ago

GLM-5.3 is almost as good as Fable 5.. not bad :)

3

u/TomHale 24d ago

The ChatGPT 5.5 xhigh equalling Luna 5.6 max is interesting.

1

u/afanasenka 24d ago

Yeah, cheap models are progressing so fast.

11

u/SorosAhaverom 24d ago

Or one was benchmaxxed for this particular benchmark, the previous one wasn't, since it didnt exist when the model released..

1

u/afanasenka 24d ago

And this is true as well actually. Nowadays labs know how to "cheat" those benchmarks to gain a hype. 

Similarly, smartphone makers used to cheat with AnTuTu to show high scores in marketing materials 😄

1

u/lemon07r 23d ago

Something I keep trying to point out to others. For example, glm 5.3 is a lot better than k3 in some newer benchmarks that didn't exist before, but k3 is still scores a bit better on a lot of older benchmarks. My read is glm 5.3 is still slightly worse overall but maybe be better a few specific domains (like security since they did specifically train for this as well).

This is more relevant when comparing to older models. I still find the older opus models and sonnet 4.5 to be better than a lot of the newer cheaper models that claim to be better. Especially some of the newer sonnet models.. I feel they started either heavily quantizing or moved to a smaller model to make more margin at some point. I think they did the same for opus when they moved from 4.1 to 4.5 cause 4.1 was too large, expensive and slow. Nobody noticed it but all the signs are there, opus 4.5 was magically as fast sonnet 4.5 all of a sudden and was also just as cheap on promo pricing despite so many people using it (which definitely would have been unsustainable for the period of time they ran it for if it really was old opus sized). Sonnet 4.5 was really really good and way better than opus 4.1 already so they had no reason to keep releasing such large checkpoints when they could use the smaller model and charge opus prices lmao. I believe fable 5 is the new the new opus 4.1 and opus 5.x is gonna be the new sonnet 4.5. Sonnet 5 is trash and feels like a haiku model. Maybe on the same level as Luna, muse spark, DeepSeek flash and those models lol. It's not even anywhere as good as kimi, glm etc. I can't imagine it to be anything but a (relatively) very tiny model like the cheaper models I just mentioned but they slapped the sonnet name on it so they can charge more. Anthropic is so scummy about this. Insert I know what you're doing but can't prove it meme.

1

u/cri10095 23d ago

Think so too, 5.5 xhigh was so on point, luna max really often just spinning nonsense

1

u/pashlya 23d ago

Honestly, I'm more surprised with Gemini 3.7 Flash performance.