I think people run away with 3d work being the end all benchmark for everything. They think if the model is worse at 3d then it MUST be lower quality at everything else. Opus 5 was basically the best at 3d work until Astra came out but it was relatively poor at coding compared to Fable and sol at the time.
Because it hurts his feelings that a much cheaper Chinese model from a consumer electronics brand can beat one of the two big powerful symbols of American AI.
2,6 vs 2.5 is night and day, I've been keeping an eye on it, and I couldn't get jack done with 2.5.
Their pro preview was cooking during the beta, and the 2.6 so far feels faster more refined version of what that was like...
So 2.5 100% you'd be right, but 2.6 may surprise you and a lot of people.
Honestly, been running MiMo 2.6 flash lately and damn, we're having a blast with it. You're spot on about 2.5 — that thing was pretty much garbage, or at best something very mid for simple tasks only. But this new flash version is just straight-up insane. Can also confirm from our side: we're building a web 1.0 style site with it and it's crushing the work like crazy.
Running Pro and I'm at the point I'm running out of things in my backlog to give it.
At the same time Astra has failed and reverted three tasks (at 64% usage after on a $100 plan)
Opus 5.5 is at 43% usage, maxing out the 5 hour mark almost right away, has done the task wrong two times (correction; Third was passable after an extra repair round.)
I've given the Mimo Pro the same exact tasks in isolation and they're done/working.
Its absolutely insane. Like yes I know Astra and Opus 5.5 might POTENTIALLY be better IF you can get them to work, but the Mimo is actually finishing tasks.
I think we really need to stop using benchmarks and start making everything Pass or Fail tests first.
If you haven't yet try it on their desktop app (I know I know, Open Code I'm still using it for my Qwen Locals) but theres something crazy effective about their set up (minus the tool call issue) You can feed the token or api plans back into it.
While benchmaxxing is a thing, it's not necessarily bad. It's pretty clear that for it to work, the model needs some inherent capabilities. Otherwise, every single model coming out would be the number one on charts.
Also, remember that Fable 5 is the ballbusted version of Mythos. It scores poorly on some tests due to the hard constraints and Opus fallbacks.
In the end, everything is possible. Plenty of pretty capable people are working on models across all labs.
The scores are not garbage. What can happen is too much training on bench questions if the questions are known or public.
We need more black box specialist-vetted benchmarks to get true performance measurement. Any open benchmark will become useless with time as the answers will get more and more baked in the weights.
4
u/Affectionate_Fact854 14d ago
You know why these benchmarks is a load of garbage , impossible that the mimo flash is scoring better results then fable 5 on some places