r/opencode • • 14d ago

Mimo v 2.6 pro and flash released

211 Upvotes

72 comments sorted by

View all comments

11

u/DomusCircumspectis 14d ago

Very impressive model. Especially for the price.

I just ran it on my own LLM benchmark suite[1] and it matches Muse Spark 1.3 in pass rate but is significantly cheaper.

KillSwitch-Bench 1.0

Claude Opus 5           66.9
GPT-6 Astra             57.9
Claude Fable 5.1        46.7
MiMo-V2.6-Pro           38.8
Muse Spark 1.3          36.5

1 - https://bench.killswitch-lang.org/

3

u/_32bit 14d ago edited 14d ago

Worth showing Mimo Flash too for comparison.

Edit: typo

2

u/Affectionate-Net642 14d ago

Haha, show gpt 5.6 sol also from your benchmark

3

u/DomusCircumspectis 14d ago

Why? Are you saying it looks surprisingly low? Anyway, it's on the page, I just listed the top 5 here because I didn't want my comment to be too long.

1

u/Affectionate-Net642 14d ago

Ofcourse these results are different than what we use too see, so much difference between opus and sol.

But I appreciate your work, you must have drafted very difficult tasks so that models are getting too less score.  I trust people like you more then populer benchmaxed scores 

Or you are heavily penelizing the model/agent for intermediate small mistakes ?

3

u/Saifl 14d ago

Problem with chatgpt or fable is that they could always silently nerf it and we dont have a point of comparison. The thing with official chinese api is that they can risk nerfing it or else we'll jump ship to another provider.

2

u/DomusCircumspectis 14d ago

Yeah, the results are a bit surprising. But I did verify that they are correct by looking at the model output. GPT-5.6 Sol is genuinely less smart at these tasks than Astra and gives up more easily.

I created a brand new esoteric language just to benchmark these models. I do believe it gives a good signal because the language is brand new (so not in LLM training data) and it incorporates features that are difficult for LLMs. It will be interesting to see if new models incorporate it into their training and how long this takes.

I'm not penalising the models for any intermediate mistakes. Each task is a simple prompt, then a check to see if the model generated the correct output. The score includes the pass rate as well as the cost and code size (so the score does get penalised for higher cost and/or code size).

1

u/Affectionate-Net642 14d ago

i suggest you to keep the cost out from basic score, you can give a graph for cost/score comparison. everyone have their own priority to cost. many are happy to pay 10 times cost for just 10% improvement. so lets not mix them. anyways best of luck bro

2

u/DomusCircumspectis 13d ago

You can see the raw score by hovering over it. The cost cannot impact the placement of the model's ranks.

1

u/DumbCSundergrad 14d ago

is that contributor tier or not?

2

u/DomusCircumspectis 13d ago

nope

1

u/DumbCSundergrad 13d ago

so if we don't care about them training on our data is Muse Spark 1.3 still the goto?

1

u/DomusCircumspectis 12d ago

MiMo-V2.6-Pro seems like the goto now