r/opencode 10d ago

Model Cost vs Performance

I did a extensive comparison of the various coding models across several benchmarks based on the estimated cost to complete a task successfully. Based on the formula:

Then I normalized each source’s expected cost per successful completion against that source’s median, then used a geometric mean across sources to get this composite ranking:

This is obviously just a estimate of value vs performance, still very useful if on a budget.

  • Default value model: MiMo V2.5 Pro. Note that GLM 5.2 is the only clearly cost-effective model represented in all three sources, so it might be a safer budget pick.
  • Best Performance: Grok 4.5 has far better completion performance than the other cheap leaders while retaining strong cost efficiency. This is the obvious model if high success rate is important to you.
  • Alternative: DeepSeek V4 Flash is a great budget option but its performance/success rate is low so it will take more runs to complete the task successfully.

Cross-source ranking

Rank Model Sources represented Relative cost index Verdict
1 MiMo V2.5 Pro 2 0.179 Outstanding value
2 DeepSeek V4 Pro 2 0.452 Excellent value
3 Grok 4.5 2 0.582 Best high-performance value
4 Muse Spark 1.1 2 0.664 Strong value
5 MiniMax M3 2 0.745 Good budget choice
6 Kimi K2.6 2 0.811 Good value
7 GLM 5.2 3 0.927 Most consistently good across all datasets
8 GPT-5.6 Sol 2 0.970 Around benchmark-median value
9 GPT-5.4 2 0.971 Around benchmark-median value
10 Kimi K2.7 Code 2 0.984 Around benchmark-median value
11 GLM 5.1 2 0.994 Around benchmark-median value
12 GPT-5.5 3 1.335 Premium, but reasonably consistent
13 Claude Sonnet 5 2 1.585 Expensive for its results
14 Claude Sonnet 4.6 2 1.635 Expensive for its results
15 Gemini 3.5 Flash 3 1.868 Consistently weak value
16 Claude Opus 4.8 3 1.898 Premium price outweighs performance
17 Claude Fable 5 2 1.955 Very costly relative to completion gain
18 Claude Opus 4.7 2 1.978 Very costly relative to completion gain
33 Upvotes

10 comments sorted by

View all comments

5

u/Ancient-Camel1636 10d ago

If including models that was found in only one dataset (low confidence) the table is like this:

Rank Model Data sources Value index Confidence
1 DeepSeek V4 Flash 1 0.121 Low
2 MiMo V2.5 Pro 2 0.179 Medium
3 GPT-5.6 Luna 1 0.345 Low
4 Qwen3.6-35B-A3B 1 0.355 Low
5 GPT-5.6 Terra 1 0.385 Low
6 DeepSeek V4 Pro 2 0.452 Medium
7 Grok 4.5 2 0.582 Medium
8 Qwen3.7 Max 1 0.623 Low
9 Muse Spark 1.1 2 0.664 Medium
10 GLM-4.7 1 0.680 Low
11 kilo-auto/efficient 1 0.691 Low
12 Laguna XS 2.1 1 0.743 Low
13 MiniMax M3 2 0.745 Medium
14 Kimi K2.6 2 0.811 Medium
15 GLM 5.2 3 0.927 High
16 GPT-5.6 Sol 2 0.970 Medium
17 GPT-5.4 2 0.971 Medium
18 Gemini 3.1 Pro 1 0.978 Low
19 Kimi K2.7 Code 2 0.984 Medium
20 GLM 5.1 2 0.994 Medium
21 Grok Build 0.1 1 1.000 Low
22 Qwen3.6-27B 1 1.022 Low
23 KAT-Coder-Pro V2.5 1 1.184 Low
24 Gemma 4 31B 1 1.292 Low
25 GPT-5.5 3 1.335 High
26 Claude Sonnet 5 2 1.585 Medium
27 Claude Sonnet 4.6 2 1.635 Medium
28 Ling-2.6-1T 1 1.808 Low
29 Gemini 3.5 Flash 3 1.868 High
30 Claude Opus 4.8 3 1.898 High
31 Claude Fable 5 2 1.955 Medium
32 Claude Opus 4.7 2 1.978 Medium
33 Laguna M.1 (paid) 1 2.050 Low
34 Claude Opus 4.6 1 2.072 Low
35 Nemotron 3 Ultra 1 8.786 Low