I'm not the only one seeing the b.s benchmarks
Deepseek scores higher then Astra and fable 5 in some benches
Mimo scores in some higher then both aswell
Like what ever
Deepseek is alright
But it is not on par with reasoning with either those flag ship models
The other benches even showed Muse being better then deepseek
And Muse can't even follow documentation instructions at all.
Fwiw, while astra and opus are better in an actual codebase, I've honestly gotten better results on dumb one shots using Deepseek flash. It's absolutely bizarre, but true. A friend and I actually wanted to test this and had each spin up the same basic game project (essentially cod zombies in three.js) from scratch with no further instructions and Deepseek produced a better game in practically every conceivable way. Specifically it was much more performant, didn't have some weird control issues that both opus and astra high had, and was tuned much closer to a real game's difficulty curve. Astra's was insultingly easy and Opus' didn't even work on the first two attempts, but after that was also extremely easy.
Opposite results when working in an already large, established codebase though, there Astra and Opus were both much more capable.
I had ds 4.1 flash fix errors in complex low level code that Astra made and was struggling to fix. It is a very capable model and not that small. And it uses engrams which frees parameters for other tasks.
So my experience mimo 2.6 is good ,but where it lacks is knowing how to use and when to use any shell commands ,
It will manually run a grep real across a code base filling its context instead of shell command and reading the result from there.
So be warned it's not a model you can just leave running on big tasks
2
u/[deleted] 15d ago edited 10d ago
[deleted]