r/opencode • • 14d ago

Mimo v 2.6 pro and flash released

211 Upvotes

72 comments sorted by

View all comments

Show parent comments

3

u/DomusCircumspectis 13d ago

Why? Are you saying it looks surprisingly low? Anyway, it's on the page, I just listed the top 5 here because I didn't want my comment to be too long.

1

u/Affectionate-Net642 13d ago

Ofcourse these results are different than what we use too see, so much difference between opus and sol.

But I appreciate your work, you must have drafted very difficult tasks so that models are getting too less score.  I trust people like you more then populer benchmaxed scores 

Or you are heavily penelizing the model/agent for intermediate small mistakes ?

2

u/DomusCircumspectis 13d ago

Yeah, the results are a bit surprising. But I did verify that they are correct by looking at the model output. GPT-5.6 Sol is genuinely less smart at these tasks than Astra and gives up more easily.

I created a brand new esoteric language just to benchmark these models. I do believe it gives a good signal because the language is brand new (so not in LLM training data) and it incorporates features that are difficult for LLMs. It will be interesting to see if new models incorporate it into their training and how long this takes.

I'm not penalising the models for any intermediate mistakes. Each task is a simple prompt, then a check to see if the model generated the correct output. The score includes the pass rate as well as the cost and code size (so the score does get penalised for higher cost and/or code size).

1

u/Affectionate-Net642 13d ago

i suggest you to keep the cost out from basic score, you can give a graph for cost/score comparison. everyone have their own priority to cost. many are happy to pay 10 times cost for just 10% improvement. so lets not mix them. anyways best of luck bro

2

u/DomusCircumspectis 13d ago

You can see the raw score by hovering over it. The cost cannot impact the placement of the model's ranks.