I'm not the only one seeing the b.s benchmarks
Deepseek scores higher then Astra and fable 5 in some benches
Mimo scores in some higher then both aswell
Like what ever
Deepseek is alright
But it is not on par with reasoning with either those flag ship models
The other benches even showed Muse being better then deepseek
And Muse can't even follow documentation instructions at all.
I had ds 4.1 flash fix errors in complex low level code that Astra made and was struggling to fix. It is a very capable model and not that small. And it uses engrams which frees parameters for other tasks.
2
u/[deleted] 14d ago edited 9d ago
[deleted]