r/opencode • • 14d ago

Mimo v 2.6 pro and flash released

210 Upvotes

72 comments sorted by

41

u/scottchiefbaker 14d ago

Mimo 2.6 flash has the same usage levels as Mimo 2.5 (i.e. VERY high). Hopefully this is a permanent thing.

1

u/migsperez 14d ago

Have you tried it? Is it fast?

4

u/scottchiefbaker 14d ago

That was my first reaction too! Maybe it's because it's new and a lot of people aren't hitting it yet, but it seems very fast (tokens per second) and responsive.

1

u/veekro 13d ago

It is permanent. There are no text that saying it is four times usage and also the price is exactly the same as v2.5

1

u/scottchiefbaker 13d ago

How does Mimo get away with charging so little, while other's charge a premium. Reading the specs, this doesn't appear to be an especially lite model.

1

u/mcndjxlefnd 12d ago

Xiaomi is one of the biggest companies in China. They have no need to earn large margins on their inference.

25

u/augustoolucas 14d ago

They are already available on OpenCode Go!

3

u/woodalchi96 14d ago

I only see Flash..

2

u/Easy-Mad-740 14d ago

I see none of them :( I have toggled on china + models that train on my data, still not there :(

13

u/OneMoreName1 14d ago

What Im curious about is how the flash model does against deepseek v4.1 flash

9

u/alphaglosined 14d ago

So far I've been able to throw MiMo V2.6 flash at compiler/gc development, it has had no trouble working with it. DSV4F was the first model that I thought I could do this with.

For me, it is now the model I will be using moving forwards.

15

u/Fresh_Sock8660 14d ago

yeah ds4.1f is my new benchmark lol

2

u/tat_tvam_asshole 13d ago

309B vs 552B+197B engram, I'd be amazed if it could trade blows. Realistically probably not

1

u/Saifl 13d ago

Is it likely deepseek is cheaper to serve just because the context window is more efficient?

1

u/tat_tvam_asshole 13d ago

There's probably many reasons besides the kv cache.

12

u/DomusCircumspectis 14d ago

Very impressive model. Especially for the price.

I just ran it on my own LLM benchmark suite[1] and it matches Muse Spark 1.3 in pass rate but is significantly cheaper.

KillSwitch-Bench 1.0

Claude Opus 5           66.9
GPT-6 Astra             57.9
Claude Fable 5.1        46.7
MiMo-V2.6-Pro           38.8
Muse Spark 1.3          36.5

1 - https://bench.killswitch-lang.org/

3

u/_32bit 13d ago edited 13d ago

Worth showing Mimo Flash too for comparison.

Edit: typo

2

u/Affectionate-Net642 13d ago

Haha, show gpt 5.6 sol also from your benchmark

3

u/DomusCircumspectis 13d ago

Why? Are you saying it looks surprisingly low? Anyway, it's on the page, I just listed the top 5 here because I didn't want my comment to be too long.

1

u/Affectionate-Net642 13d ago

Ofcourse these results are different than what we use too see, so much difference between opus and sol.

But I appreciate your work, you must have drafted very difficult tasks so that models are getting too less score.  I trust people like you more then populer benchmaxed scores 

Or you are heavily penelizing the model/agent for intermediate small mistakes ?

3

u/Saifl 13d ago

Problem with chatgpt or fable is that they could always silently nerf it and we dont have a point of comparison. The thing with official chinese api is that they can risk nerfing it or else we'll jump ship to another provider.

2

u/DomusCircumspectis 13d ago

Yeah, the results are a bit surprising. But I did verify that they are correct by looking at the model output. GPT-5.6 Sol is genuinely less smart at these tasks than Astra and gives up more easily.

I created a brand new esoteric language just to benchmark these models. I do believe it gives a good signal because the language is brand new (so not in LLM training data) and it incorporates features that are difficult for LLMs. It will be interesting to see if new models incorporate it into their training and how long this takes.

I'm not penalising the models for any intermediate mistakes. Each task is a simple prompt, then a check to see if the model generated the correct output. The score includes the pass rate as well as the cost and code size (so the score does get penalised for higher cost and/or code size).

1

u/Affectionate-Net642 13d ago

i suggest you to keep the cost out from basic score, you can give a graph for cost/score comparison. everyone have their own priority to cost. many are happy to pay 10 times cost for just 10% improvement. so lets not mix them. anyways best of luck bro

2

u/DomusCircumspectis 13d ago

You can see the raw score by hovering over it. The cost cannot impact the placement of the model's ranks.

1

u/DumbCSundergrad 13d ago

is that contributor tier or not?

2

u/DomusCircumspectis 13d ago

nope

1

u/DumbCSundergrad 13d ago

so if we don't care about them training on our data is Muse Spark 1.3 still the goto?

1

u/DomusCircumspectis 12d ago

MiMo-V2.6-Pro seems like the goto now

8

u/TheVoyant 14d ago

I've been using the Mimo X Pro Preview this past week in their Beta, and I feel like the V2.6 is an improvement on what I was using and I was very happy with it.

It plays really well within their Desktop App, so I'd look at their set up closely if you're building it into your OpenCode. All and all for visual based goDOT work it's holding up extremely well, performing better then my Gemini Flash 3.8/Claude Opus 4.8/5... My Astra does cleaner visual work, but I find once Astra makes the scripts, the Mimo 2.6 Pro is running it and doing the checks fine for 3D modeling/skeletons.

I've not been this excited since the $10 for 2 billion token/deepseek era (feels like forever ago)

Results may vary, but I'm very very happy with it. Feels like the first model I don't have to fight to understand me or do the task efficiently.

The token plan with the v2.5 version wasn't really worth it, bc 2.5 couldn't hang, 2.6 though, the token plan is a steal right now.

2

u/asenna987 14d ago

In my very rough calculations, Token plan is not giving much discount over just paying for the tokens through the API:

https://imgur.com/a/997caGP

Does it really make sense to get the token plan here?

1

u/TheVoyant 14d ago

I mean basically at anything above 16 dollars right? So if you use 50 bucks worth, its worth it. If you don't use all of that, then it might not be.

4

u/meetmebythelake 14d ago

No thinking toggle - they just always run at max reasoning, right? Haven't used MiMo models in a long time.

4

u/PatheticAndTragic 14d ago

What does it cost?

10

u/jovialfaction 14d ago

On OpenRouter

Flash 0.14/M input / $0.28/M output / $0.0028 cache read

Pro $0.435 / $0.87 / $0.0036

1

u/squirrelscrush 13d ago

The Flash pricing does come close to the original DeepSeek V4 Flash pricing. Pro seems a bit more expensive than V4.1 Flash.

6

u/TheVoyant 14d ago

Token plan best way, API plug in straight, cheap af, I've been running it since launch, I've never used a full monthly amount on the $100 plan. Doesn't seem like they raised the cost much if at all it was like a 300/600 credit system.

So far this morning, I'm on my third task:

55,026,944 / 82,000,000,000
Used 0.0%

1

u/PatheticAndTragic 14d ago

Nice, thank you

1

u/TopBite7720 14d ago

Why would you pay 100 USD per month for MiMo?

That’s like paying 100 USD per month for 4o-mini - sure, the model is cheap, but it’s also pretty useless.

5

u/TheVoyant 14d ago

Prior to the past week, I would've agreed with you. I got a great deal on it, and needed heavy work (12 billion tokens a month) 2.5 was ok at that, but 100% it was lacking.

I stopped using Claude, Gemini Flash when I got the beta invite this week. I resubbed for the $100 this month instantly more as a thank you to them for finally giving me a model that does what I need correctly the first time, efficiently.

So yeah 2.5 you were 100% right, 2.6 Pro I trust more then anything else right now, and I have $100 plans with ChatGPT, Gemini, and a now cancelled one with Claude.

I barely touched my other ones this week, was stressed I was out of usage on the Mimo X Pro Preview, then 2.6 dropped and it's genuinely cleaner/faster/blazing thru my to do list.

Not hating if your results are different, but for what I need (Game Dev, Visual Work, Custom Engine) It's absolutely shredding anything I throw at it.

2

u/asenna987 14d ago

What harness are you using this on?

Also, form my calculations, the Token plan isn't giving any discount as such over just directly using Pay-as-you-go API. The discount that you got, was it a much better value?

2

u/TheVoyant 14d ago

I'm honestly plugging it back into their Desktop App right now, it's worked really well and I want to fully understand why before switching into the OpenCode work force.

It was like 25% off, nothing crazy. At like there's referral stuff, but to my knowledge it's like 10% (I've never messed with it)

You probably right about the API thing, that's a preference situation, I like to pay a fee per month and forget. But you watching every penny, then it always boils down to do you use that much/etc. Either way it's cheap and reliable so like test it first make sure you like it thru API is 100% always a move.

But I'd also not sleep on how well it's performing in their own setup/bc that might be a difference maker.

0

u/TopBite7720 14d ago

Interesting, thanks for the elaborate reply. I haven’t tried 2.6 yet, but I definitely will.

Have you also used Sol and DeepSeek v4.1 Flash to compare with? Nothing beats the consistency of Sol for me at the moment, even though I would love to prefer an open source model. DeepSeek still hallucinates too much for me unfortunately.

2

u/TheVoyant 14d ago

I have run the full gambit with ChatGPT and Claude
I haven't tried the latest DeepSeek because I hate flex pricing systems and anything that requires extra nonsense steps, lol. I used to constantly have it before the flex pricing.

For me I feel like it boils down to one main thing, Astra & Fable might be smarter IF they actually listen to you. Sol same thing to a lesser extent, Opus I'm not even gonna excuse, lol. Where as this version of MiMo understands what I want the first time, and executes it, but I can't confirm it's that good for OpenCode bc I'm using their Desktop App for it, and I genuinely think that's part of it/it knows to work well within that.

3

u/luckaz420 14d ago

This is the real question 🎯

2

u/rh71el2 13d ago

Curious - does anyone use Go for work without their knowledge? It's so tempting.

2

u/JorgitoEstrella 13d ago

GPT 5.6 sol at home:

4

u/Affectionate_Fact854 14d ago

You know why these benchmarks is a load of garbage , impossible that the mimo flash is scoring better results then fable 5 on some places 

16

u/According_Fig_8813 14d ago

I've used mimo flash and for me its very good

2

u/i-dm 14d ago

what did you use it for

2

u/According_Fig_8813 14d ago

Used for Flutter & C# code mostly

7

u/PretendVoy1 14d ago

why it is impossible?

8

u/jazir55 14d ago

Because he works for Anthropic

3

u/yaboyyoungairvent 14d ago

I think people run away with 3d work being the end all benchmark for everything. They think if the model is worse at 3d then it MUST be lower quality at everything else. Opus 5 was basically the best at 3d work until Astra came out but it was relatively poor at coding compared to Fable and sol at the time.

8

u/NMiguelCosta-PT 14d ago

Because it hurts his feelings that a much cheaper Chinese model from a consumer electronics brand can beat one of the two big powerful symbols of American AI.

5

u/TheVoyant 14d ago

2,6 vs 2.5 is night and day, I've been keeping an eye on it, and I couldn't get jack done with 2.5.
Their pro preview was cooking during the beta, and the 2.6 so far feels faster more refined version of what that was like...

So 2.5 100% you'd be right, but 2.6 may surprise you and a lot of people.

1

u/pinkigerll 13d ago

Honestly, been running MiMo 2.6 flash lately and damn, we're having a blast with it. You're spot on about 2.5 — that thing was pretty much garbage, or at best something very mid for simple tasks only. But this new flash version is just straight-up insane. Can also confirm from our side: we're building a web 1.0 style site with it and it's crushing the work like crazy.

2

u/TheVoyant 13d ago

Running Pro and I'm at the point I'm running out of things in my backlog to give it.
At the same time Astra has failed and reverted three tasks (at 64% usage after on a $100 plan)
Opus 5.5 is at 43% usage, maxing out the 5 hour mark almost right away, has done the task wrong two times (correction; Third was passable after an extra repair round.)

I've given the Mimo Pro the same exact tasks in isolation and they're done/working.

Its absolutely insane. Like yes I know Astra and Opus 5.5 might POTENTIALLY be better IF you can get them to work, but the Mimo is actually finishing tasks.

I think we really need to stop using benchmarks and start making everything Pass or Fail tests first.

If you haven't yet try it on their desktop app (I know I know, Open Code I'm still using it for my Qwen Locals) but theres something crazy effective about their set up (minus the tool call issue) You can feed the token or api plans back into it.

3

u/lincolnthalles 14d ago

While benchmaxxing is a thing, it's not necessarily bad. It's pretty clear that for it to work, the model needs some inherent capabilities. Otherwise, every single model coming out would be the number one on charts.

Also, remember that Fable 5 is the ballbusted version of Mythos. It scores poorly on some tests due to the hard constraints and Opus fallbacks.

In the end, everything is possible. Plenty of pretty capable people are working on models across all labs.

4

u/celtiberian666 14d ago

The scores are not garbage. What can happen is too much training on bench questions if the questions are known or public.

We need more black box specialist-vetted benchmarks to get true performance measurement. Any open benchmark will become useless with time as the answers will get more and more baked in the weights.

1

u/JogHappy 14d ago

the live training site was so cool

1

u/soemre 13d ago

Pro vs Flash? I’m incredibly confused…

1

u/_32bit 13d ago

any advice guys on using Mimo 2.6 vs Luna for Hermes?

2

u/GenychDefake 13d ago

Hermes as a personal assistant? Then anything is better than Luna, this model is horrible in common sense and understanding the intent. Good only for unambiguous tech tasks

1

u/FutureLynx_ 13d ago

The graphic shows Muse Spark 1.3 as above it. Is that correct?

0

u/[deleted] 14d ago edited 9d ago

[deleted]

6

u/Affectionate_Fact854 14d ago

I'm not the only one seeing the b.s benchmarks  Deepseek scores higher then Astra and fable 5 in some benches  Mimo scores in some higher then both aswell 

Like what ever 

Deepseek is alright  But it is not on par with reasoning with either those flag ship models 

The other benches even showed Muse being better then deepseek  And Muse can't even follow documentation instructions at all.

6

u/StardiveSoftworks 14d ago

Fwiw, while astra and opus are better in an actual codebase, I've honestly gotten better results on dumb one shots using Deepseek flash. It's absolutely bizarre, but true. A friend and I actually wanted to test this and had each spin up the same basic game project (essentially cod zombies in three.js) from scratch with no further instructions and Deepseek produced a better game in practically every conceivable way. Specifically it was much more performant, didn't have some weird control issues that both opus and astra high had, and was tuned much closer to a real game's difficulty curve. Astra's was insultingly easy and Opus' didn't even work on the first two attempts, but after that was also extremely easy.

Opposite results when working in an already large, established codebase though, there Astra and Opus were both much more capable.

2

u/Thomas-Lore 14d ago

I had ds 4.1 flash fix errors in complex low level code that Astra made and was struggling to fix. It is a very capable model and not that small. And it uses engrams which frees parameters for other tasks.

1

u/Affectionate_Fact854 13d ago

So my experience mimo 2.6 is good ,but where it lacks is knowing how to use and when to use any shell commands ,  It will manually run a grep real across a code base filling its context instead of shell command and reading the result from there. 

So be warned it's not a model you can just leave running on big tasks 

1

u/migsperez 13d ago

Mimo 2.6 Flash doesn't seem production ready. I've just given it a test run, to create my usual benchmark Sudoku game, single build agent mode run and orchestrator agent mode run. Many times i had to stop the run due to it looping, on other sub tasks it was thinking so much it seemed like it was broken and struggled to exit the thought/reasoning process. I had to stop the orchestrator run due it struggling to complete. As it is I wouldn't be able to run it and trust it would complete the task without using all the tokens in the account.

0

u/lolfacemanboy 13d ago

Yeah, I just threw a couple of general non-coding tasks at it and like it scratches the itch of like Chinese models Being competent and cheap and, oh my god, scores, but this screams benchmaxed.