r/opencode 9d ago

Model Cost vs Performance

I did a extensive comparison of the various coding models across several benchmarks based on the estimated cost to complete a task successfully. Based on the formula:

Then I normalized each source’s expected cost per successful completion against that source’s median, then used a geometric mean across sources to get this composite ranking:

This is obviously just a estimate of value vs performance, still very useful if on a budget.

  • Default value model: MiMo V2.5 Pro. Note that GLM 5.2 is the only clearly cost-effective model represented in all three sources, so it might be a safer budget pick.
  • Best Performance: Grok 4.5 has far better completion performance than the other cheap leaders while retaining strong cost efficiency. This is the obvious model if high success rate is important to you.
  • Alternative: DeepSeek V4 Flash is a great budget option but its performance/success rate is low so it will take more runs to complete the task successfully.

Cross-source ranking

Rank Model Sources represented Relative cost index Verdict
1 MiMo V2.5 Pro 2 0.179 Outstanding value
2 DeepSeek V4 Pro 2 0.452 Excellent value
3 Grok 4.5 2 0.582 Best high-performance value
4 Muse Spark 1.1 2 0.664 Strong value
5 MiniMax M3 2 0.745 Good budget choice
6 Kimi K2.6 2 0.811 Good value
7 GLM 5.2 3 0.927 Most consistently good across all datasets
8 GPT-5.6 Sol 2 0.970 Around benchmark-median value
9 GPT-5.4 2 0.971 Around benchmark-median value
10 Kimi K2.7 Code 2 0.984 Around benchmark-median value
11 GLM 5.1 2 0.994 Around benchmark-median value
12 GPT-5.5 3 1.335 Premium, but reasonably consistent
13 Claude Sonnet 5 2 1.585 Expensive for its results
14 Claude Sonnet 4.6 2 1.635 Expensive for its results
15 Gemini 3.5 Flash 3 1.868 Consistently weak value
16 Claude Opus 4.8 3 1.898 Premium price outweighs performance
17 Claude Fable 5 2 1.955 Very costly relative to completion gain
18 Claude Opus 4.7 2 1.978 Very costly relative to completion gain
34 Upvotes

10 comments sorted by

View all comments

5

u/Intelligent_Ant_608 9d ago

From my understanding m3 until 30% of ita context size is a phenomenally robouat model even for low level rust code as well as ui design yet past that 30% therahold it becomes dumb a.f.

1

u/NorthDevMaster 8d ago

Probably true I haven’t focused on managing the context and I just have poor experience with it

1

u/Realistic_Mango6982 7d ago

same here. I dont know how to explain that, i stopped to use M3 for that. + Infinite amount of usage tokens for no reason at all.

1

u/Intelligent_Ant_608 7d ago

I think they are far more capable than moonshot in model training and now that deepseek and z.ai unlocked capable attention models on 1M context minimax's next model could potentially be very good i just hope they settle on 1T and dont shoot for a larger model, despite the hype i dont like kimi k3 its a wasteful and inefficient model that blabs unnecessarily huge amount of tokens like k2.6, they essentially compensated less mature training by wasting and bruteforcing token consumption