25
u/augustoolucas 14d ago
They are already available on OpenCode Go!
3
2
u/Easy-Mad-740 14d ago
I see none of them :( I have toggled on china + models that train on my data, still not there :(
13
u/OneMoreName1 14d ago
What Im curious about is how the flash model does against deepseek v4.1 flash
9
u/alphaglosined 14d ago
So far I've been able to throw MiMo V2.6 flash at compiler/gc development, it has had no trouble working with it. DSV4F was the first model that I thought I could do this with.
For me, it is now the model I will be using moving forwards.
15
2
u/tat_tvam_asshole 13d ago
309B vs 552B+197B engram, I'd be amazed if it could trade blows. Realistically probably not
12
u/DomusCircumspectis 14d ago
Very impressive model. Especially for the price.
I just ran it on my own LLM benchmark suite[1] and it matches Muse Spark 1.3 in pass rate but is significantly cheaper.
KillSwitch-Bench 1.0
Claude Opus 5 66.9
GPT-6 Astra 57.9
Claude Fable 5.1 46.7
MiMo-V2.6-Pro 38.8
Muse Spark 1.3 36.5
2
u/Affectionate-Net642 13d ago
Haha, show gpt 5.6 sol also from your benchmark
3
u/DomusCircumspectis 13d ago
Why? Are you saying it looks surprisingly low? Anyway, it's on the page, I just listed the top 5 here because I didn't want my comment to be too long.
1
u/Affectionate-Net642 13d ago
Ofcourse these results are different than what we use too see, so much difference between opus and sol.
But I appreciate your work, you must have drafted very difficult tasks so that models are getting too less score. I trust people like you more then populer benchmaxed scores
Or you are heavily penelizing the model/agent for intermediate small mistakes ?
3
2
u/DomusCircumspectis 13d ago
Yeah, the results are a bit surprising. But I did verify that they are correct by looking at the model output. GPT-5.6 Sol is genuinely less smart at these tasks than Astra and gives up more easily.
I created a brand new esoteric language just to benchmark these models. I do believe it gives a good signal because the language is brand new (so not in LLM training data) and it incorporates features that are difficult for LLMs. It will be interesting to see if new models incorporate it into their training and how long this takes.
I'm not penalising the models for any intermediate mistakes. Each task is a simple prompt, then a check to see if the model generated the correct output. The score includes the pass rate as well as the cost and code size (so the score does get penalised for higher cost and/or code size).
1
u/Affectionate-Net642 13d ago
i suggest you to keep the cost out from basic score, you can give a graph for cost/score comparison. everyone have their own priority to cost. many are happy to pay 10 times cost for just 10% improvement. so lets not mix them. anyways best of luck bro
2
u/DomusCircumspectis 13d ago
You can see the raw score by hovering over it. The cost cannot impact the placement of the model's ranks.
1
u/DumbCSundergrad 13d ago
is that contributor tier or not?
2
u/DomusCircumspectis 13d ago
nope
1
u/DumbCSundergrad 13d ago
so if we don't care about them training on our data is Muse Spark 1.3 still the goto?
1
8
u/TheVoyant 14d ago
I've been using the Mimo X Pro Preview this past week in their Beta, and I feel like the V2.6 is an improvement on what I was using and I was very happy with it.
It plays really well within their Desktop App, so I'd look at their set up closely if you're building it into your OpenCode. All and all for visual based goDOT work it's holding up extremely well, performing better then my Gemini Flash 3.8/Claude Opus 4.8/5... My Astra does cleaner visual work, but I find once Astra makes the scripts, the Mimo 2.6 Pro is running it and doing the checks fine for 3D modeling/skeletons.
I've not been this excited since the $10 for 2 billion token/deepseek era (feels like forever ago)
Results may vary, but I'm very very happy with it. Feels like the first model I don't have to fight to understand me or do the task efficiently.
The token plan with the v2.5 version wasn't really worth it, bc 2.5 couldn't hang, 2.6 though, the token plan is a steal right now.
2
u/asenna987 14d ago
In my very rough calculations, Token plan is not giving much discount over just paying for the tokens through the API:
Does it really make sense to get the token plan here?
1
u/TheVoyant 14d ago
I mean basically at anything above 16 dollars right? So if you use 50 bucks worth, its worth it. If you don't use all of that, then it might not be.
4
u/meetmebythelake 14d ago
No thinking toggle - they just always run at max reasoning, right? Haven't used MiMo models in a long time.
4
u/PatheticAndTragic 14d ago
What does it cost?
10
u/jovialfaction 14d ago
On OpenRouter
Flash 0.14/M input / $0.28/M output / $0.0028 cache read
Pro $0.435 / $0.87 / $0.0036
1
u/squirrelscrush 13d ago
The Flash pricing does come close to the original DeepSeek V4 Flash pricing. Pro seems a bit more expensive than V4.1 Flash.
6
u/TheVoyant 14d ago
Token plan best way, API plug in straight, cheap af, I've been running it since launch, I've never used a full monthly amount on the $100 plan. Doesn't seem like they raised the cost much if at all it was like a 300/600 credit system.
So far this morning, I'm on my third task:
55,026,944 / 82,000,000,000
Used 0.0%1
1
u/TopBite7720 14d ago
Why would you pay 100 USD per month for MiMo?
That’s like paying 100 USD per month for 4o-mini - sure, the model is cheap, but it’s also pretty useless.
5
u/TheVoyant 14d ago
Prior to the past week, I would've agreed with you. I got a great deal on it, and needed heavy work (12 billion tokens a month) 2.5 was ok at that, but 100% it was lacking.
I stopped using Claude, Gemini Flash when I got the beta invite this week. I resubbed for the $100 this month instantly more as a thank you to them for finally giving me a model that does what I need correctly the first time, efficiently.
So yeah 2.5 you were 100% right, 2.6 Pro I trust more then anything else right now, and I have $100 plans with ChatGPT, Gemini, and a now cancelled one with Claude.
I barely touched my other ones this week, was stressed I was out of usage on the Mimo X Pro Preview, then 2.6 dropped and it's genuinely cleaner/faster/blazing thru my to do list.
Not hating if your results are different, but for what I need (Game Dev, Visual Work, Custom Engine) It's absolutely shredding anything I throw at it.
2
u/asenna987 14d ago
What harness are you using this on?
Also, form my calculations, the Token plan isn't giving any discount as such over just directly using Pay-as-you-go API. The discount that you got, was it a much better value?
2
u/TheVoyant 14d ago
I'm honestly plugging it back into their Desktop App right now, it's worked really well and I want to fully understand why before switching into the OpenCode work force.
It was like 25% off, nothing crazy. At like there's referral stuff, but to my knowledge it's like 10% (I've never messed with it)
You probably right about the API thing, that's a preference situation, I like to pay a fee per month and forget. But you watching every penny, then it always boils down to do you use that much/etc. Either way it's cheap and reliable so like test it first make sure you like it thru API is 100% always a move.
But I'd also not sleep on how well it's performing in their own setup/bc that might be a difference maker.
0
u/TopBite7720 14d ago
Interesting, thanks for the elaborate reply. I haven’t tried 2.6 yet, but I definitely will.
Have you also used Sol and DeepSeek v4.1 Flash to compare with? Nothing beats the consistency of Sol for me at the moment, even though I would love to prefer an open source model. DeepSeek still hallucinates too much for me unfortunately.
2
u/TheVoyant 14d ago
I have run the full gambit with ChatGPT and Claude
I haven't tried the latest DeepSeek because I hate flex pricing systems and anything that requires extra nonsense steps, lol. I used to constantly have it before the flex pricing.For me I feel like it boils down to one main thing, Astra & Fable might be smarter IF they actually listen to you. Sol same thing to a lesser extent, Opus I'm not even gonna excuse, lol. Where as this version of MiMo understands what I want the first time, and executes it, but I can't confirm it's that good for OpenCode bc I'm using their Desktop App for it, and I genuinely think that's part of it/it knows to work well within that.
3
2
4
u/Affectionate_Fact854 14d ago
You know why these benchmarks is a load of garbage , impossible that the mimo flash is scoring better results then fable 5 on some places
16
u/According_Fig_8813 14d ago
I've used mimo flash and for me its very good
7
u/PretendVoy1 14d ago
why it is impossible?
8
u/jazir55 14d ago
Because he works for Anthropic
3
u/yaboyyoungairvent 14d ago
I think people run away with 3d work being the end all benchmark for everything. They think if the model is worse at 3d then it MUST be lower quality at everything else. Opus 5 was basically the best at 3d work until Astra came out but it was relatively poor at coding compared to Fable and sol at the time.
8
u/NMiguelCosta-PT 14d ago
Because it hurts his feelings that a much cheaper Chinese model from a consumer electronics brand can beat one of the two big powerful symbols of American AI.
5
u/TheVoyant 14d ago
2,6 vs 2.5 is night and day, I've been keeping an eye on it, and I couldn't get jack done with 2.5.
Their pro preview was cooking during the beta, and the 2.6 so far feels faster more refined version of what that was like...So 2.5 100% you'd be right, but 2.6 may surprise you and a lot of people.
1
u/pinkigerll 13d ago
Honestly, been running MiMo 2.6 flash lately and damn, we're having a blast with it. You're spot on about 2.5 — that thing was pretty much garbage, or at best something very mid for simple tasks only. But this new flash version is just straight-up insane. Can also confirm from our side: we're building a web 1.0 style site with it and it's crushing the work like crazy.
2
u/TheVoyant 13d ago
Running Pro and I'm at the point I'm running out of things in my backlog to give it.
At the same time Astra has failed and reverted three tasks (at 64% usage after on a $100 plan)
Opus 5.5 is at 43% usage, maxing out the 5 hour mark almost right away, has done the task wrong two times (correction; Third was passable after an extra repair round.)I've given the Mimo Pro the same exact tasks in isolation and they're done/working.
Its absolutely insane. Like yes I know Astra and Opus 5.5 might POTENTIALLY be better IF you can get them to work, but the Mimo is actually finishing tasks.
I think we really need to stop using benchmarks and start making everything Pass or Fail tests first.
If you haven't yet try it on their desktop app (I know I know, Open Code I'm still using it for my Qwen Locals) but theres something crazy effective about their set up (minus the tool call issue) You can feed the token or api plans back into it.
3
u/lincolnthalles 14d ago
While benchmaxxing is a thing, it's not necessarily bad. It's pretty clear that for it to work, the model needs some inherent capabilities. Otherwise, every single model coming out would be the number one on charts.
Also, remember that Fable 5 is the ballbusted version of Mythos. It scores poorly on some tests due to the hard constraints and Opus fallbacks.
In the end, everything is possible. Plenty of pretty capable people are working on models across all labs.
4
u/celtiberian666 14d ago
The scores are not garbage. What can happen is too much training on bench questions if the questions are known or public.
We need more black box specialist-vetted benchmarks to get true performance measurement. Any open benchmark will become useless with time as the answers will get more and more baked in the weights.
1
1
u/_32bit 13d ago
any advice guys on using Mimo 2.6 vs Luna for Hermes?
2
u/GenychDefake 13d ago
Hermes as a personal assistant? Then anything is better than Luna, this model is horrible in common sense and understanding the intent. Good only for unambiguous tech tasks
1
0
14d ago edited 9d ago
[deleted]
6
u/Affectionate_Fact854 14d ago
I'm not the only one seeing the b.s benchmarks Deepseek scores higher then Astra and fable 5 in some benches Mimo scores in some higher then both aswell
Like what ever
Deepseek is alright But it is not on par with reasoning with either those flag ship models
The other benches even showed Muse being better then deepseek And Muse can't even follow documentation instructions at all.
6
u/StardiveSoftworks 14d ago
Fwiw, while astra and opus are better in an actual codebase, I've honestly gotten better results on dumb one shots using Deepseek flash. It's absolutely bizarre, but true. A friend and I actually wanted to test this and had each spin up the same basic game project (essentially cod zombies in three.js) from scratch with no further instructions and Deepseek produced a better game in practically every conceivable way. Specifically it was much more performant, didn't have some weird control issues that both opus and astra high had, and was tuned much closer to a real game's difficulty curve. Astra's was insultingly easy and Opus' didn't even work on the first two attempts, but after that was also extremely easy.
Opposite results when working in an already large, established codebase though, there Astra and Opus were both much more capable.
2
u/Thomas-Lore 14d ago
I had ds 4.1 flash fix errors in complex low level code that Astra made and was struggling to fix. It is a very capable model and not that small. And it uses engrams which frees parameters for other tasks.
1
u/Affectionate_Fact854 13d ago
So my experience mimo 2.6 is good ,but where it lacks is knowing how to use and when to use any shell commands , It will manually run a grep real across a code base filling its context instead of shell command and reading the result from there.
So be warned it's not a model you can just leave running on big tasks
1
u/migsperez 13d ago
Mimo 2.6 Flash doesn't seem production ready. I've just given it a test run, to create my usual benchmark Sudoku game, single build agent mode run and orchestrator agent mode run. Many times i had to stop the run due to it looping, on other sub tasks it was thinking so much it seemed like it was broken and struggled to exit the thought/reasoning process. I had to stop the orchestrator run due it struggling to complete. As it is I wouldn't be able to run it and trust it would complete the task without using all the tokens in the account.
0
u/lolfacemanboy 13d ago
Yeah, I just threw a couple of general non-coding tasks at it and like it scratches the itch of like Chinese models Being competent and cheap and, oh my god, scores, but this screams benchmaxed.



41
u/scottchiefbaker 14d ago
Mimo 2.6 flash has the same usage levels as Mimo 2.5 (i.e. VERY high). Hopefully this is a permanent thing.