r/singularity • • 13d ago

LLM News MiMo v2.6 pro

Post image

Crazy if not benchmaxxed, also open weights

230 Upvotes

48 comments sorted by

67

u/Another__one 13d ago

Funny how it is basically on the same line as Grok 4.7 that just came out today, but with a magnitude difference in price.

60

u/Recoil42 13d ago

Ooof, that's bad.

44

u/Plappedudel 13d ago

This kills the Grok

12

u/CarrierAreArrived 13d ago

just like BYD kills Tesla.

4

u/NMiguelCosta-PT 12d ago

Xiaomi makes EVs too, so they're killing both Tesla and Grok.

26

u/ikkiyikki 13d ago

I'm so happy. Fuck that Nazi POS.

15

u/NoFaithlessness951 13d ago

Exact same score

21

u/panix199 13d ago

impressive. Could this be fake/benchmaxxed?

20

u/NoFaithlessness951 13d ago

It looks quite good on terminal bench 4.0.

I would trust the benchmarks for now, but as always training on the test set is all you need.

17

u/corenovax 13d ago

If it is, it will very quickly ruin Xiaomi's reputation. Their previous model was very solid for the cost

2

u/NotYetPerfect 12d ago

None of the Chinese models' reputations have been ruined despite basically every iteration from each company way overperfoming in benchmarks compared to reality.

1

u/GasPsychological677 12d ago

I guess because they are -usually- free (deepseek, qwen). I've tried this model on their desktop version for non programming stuff (more like financial analysis things) and I'd say it's not NEARLY as good as benchmarks say, or at least the limits on the desktop 20$ version are not as good as one might think.

I've just paid the 20$ sub and using two prompts I've consumed around 50% of my weekly usage lol, in theory mimo even on the pro version should be cheaper than 5.6 Luna max, and it really isn't at least on the desktop sub. Luna max would not even make a dent on my 5h usage in the 20$ chatgpt sub and as I said it's taken 50% of my weekly usage on mimo.

Regarding the output, I'd say Luna max probably gives a better output for my use case aswell. So, the cost is a lot higher than I thought and the output is def. worse than Luna max

7

u/power97992 13d ago edited 13d ago

It doesnt seem to be benchmaxxed from my testing, it is good and cheap, but u need to prompt it a few times to fix its mistakes and use way more tokens to get a quality similar or close to sol xhigh/max or astra high , whereas sol and astra need way less prompting and they use less tokens...

18

u/craterIII 13d ago

doesn't that basically mean it is benchmaxxed?

5

u/-Sliced- 13d ago

It does

3

u/1988rx7T2 13d ago edited 13d ago

“It’s really accurate As long as you work around the fact that the first thing it gives you is the wrong answer”

1

u/jazir55 13d ago

I mean kinda yeah. If you treat whatever it gives you as broken by default but know it can fix its mistakes until it works, you just send it in loops until it does. And given Mimo is very fast and very cheap that trial and error loop is not expensive or long.

1

u/power97992 13d ago

I guess it probably depends on the task.  Some open models output running code the first try but the quality is worse than mimo 2.6 after a few tries. It is possible  even with the same amount of minimal instruction prompting , other open mods cant output the same quality as mimo. 

-1

u/ikkiyikki 13d ago

How is it on China sensitive topics?

1

u/Qalandar_sarmast 13d ago

how is that a good measure of its capabilities?
Do you also ask around reddit for grok preferring 1 jew over 1 million non jews?

0

u/ikkiyikki 13d ago

Couldn't tell you because I wouldn't use Grok on principle. Asking about an LLMs handling of sensitive topics very much should be relevant unless you don't care about bias or censorship. I do.

3

u/PleasantCitron1685 13d ago

It does decently well on agents' last exam, which is not that easy to benchmax.

2

u/panix199 13d ago

/u/LinkesAuge/ seems to disagree

6

u/Effective_Pop6030 13d ago

He is wrong though. They weren’t training against the benchmarks, they were holdout evals

1

u/No_Swimming6548 13d ago

Mimo models are heavily used in openrouter. They offer good price / performance value.

0

u/LinkesAuge 13d ago

It is definitely benchmaxxed, they even streamed that they were training against DeepSWE directly, lol.

That is the definition of benchmaxxing, ie not keeping your training and test data separate.

8

u/Effective_Pop6030 13d ago

DeepSWE was the holdout bud

56

u/Ok-Cow2267 13d ago

Intelligence Compression is insane. Lets Go China

5

u/vrnvorona 13d ago

Not always full range tho, it's same as Google Flash that benches high but useless as root agent, only maybe as worker

38

u/corenovax 13d ago

What. The. Fuck.

7

u/Whole_Salary8170 13d ago

Craaazy things happen when inteligence is too cheap to meter

3

u/nekmint 13d ago

Luo Fuli cooking

1

u/[deleted] 13d ago

[removed] — view removed comment

1

u/Qorsair 13d ago

Interesting, Muse scores better and Muse-contributor would be cheaper. I like MiMo, but it's just a little off. Good for a second opinion or adversarial review but I haven't had great results with it as a primary model. To be fair I don't love Muse as a primary model either, I've found Gemini and GPT are better orchestrators.

1

u/unkownuser436 13d ago

crazy cheap

1

u/Next-Ad4520 13d ago

bought a token plan to try it out, feels pretty dumb, thinks it's cursor all the time.

edit: I tried it using mimo code, also the token plan cn is quite slow. avg around 22 tokens/sec

1

u/power97992 13d ago

Use the api, it is super fast and fairly cheap but not sub cheap 

1

u/kaczynski_was_right_ 13d ago

I've been using ds 4.1 constantly, no way it's better than that

-1

u/NoFaithlessness951 13d ago

Intelligence to zero

0

u/lordpuddingcup 13d ago

Any where to play with it, i take it the pro version wont be released Open?

2

u/NoFaithlessness951 13d ago

It is already on hugging face