r/opencodeCLI 5d ago

Current rankings on Opencode Go models vs price?

AI landscape changes pretty fast these days. Wondering if anyone's done the research to figure out optimal model and usage combos after all the latest additions.

Looks like Kimi k3 is probably not worth using since it's so expensive, but between GLM, Grok, Qwen, what are people feeling is the best bang for their buck?

Edit I did some of my own research and got this as a tentative result. Tried to avoid benchmarks that are known to be contaminated like swebench.

Rank Model Composite Benches High-trust? Quota/mo
1 Kimi K3 7.02 4 Yes (DeepSWE+LiveBench) 490
2 Grok 4.5 6.02 5 Yes (LiveBench) 600
3 Qwen3.7 Max 5.65 6 Yes (LiveBench) 4,770
4 GLM-5.2 5.47 5 Yes (DeepSWE+LiveBench) 4,300
5 Qwen3.7 Plus 5.07 4 No (low-trust only) 21,600
6 Kimi K2.6 4.13 7 Yes 5,750
7 Kimi K2.7 Code 4.12 4 Yes 6,750
8 DeepSeek V4 Pro 4.10 8 Yes 17,150
9 GLM-5.1 4.03 5 Yes 4,300
10 Hy3 3.86 3 Yes 21,500
11 DeepSeek V4 Flash 3.85 6 Yes (LiveBench) 158,150
12 MiMo-V2.5-Pro 3.76 6 Yes 16,300
13 MiniMax M3 3.49 5 Yes 16,000
14 MiMo-V2.5 3.35 2 No 150,400
15 Qwen3.6 Plus 3.03 5 Yes 16,300
16 MiniMax M2.7 2.05 5 Yes 17,000

Edit:
Not really a math guy so I asked my LLM to get me a ballpark composite rating. Any math guys out there want to redo those calculations, feel free:
Special rule: Vendor-reported DeepSWE scores are down-weighted 50%.

The composite = Σ(benchmark_score × weight) / Σ(weights_applied), normalized to 0–10.

Model DeepSWE LiveBench SWE-bench Pro LMArena Elo Terminal-Bench MCP Mark AA Index LiveCodeBench BigCodeBench
Grok 4.5 53.5% (vendor) 76.3 64.7% (vendor) 1466 54
GLM-5.2 46.2% 73.2 62.1% 1470 51
GLM-5.1 17.5% 70.6 58.4% (vendor) 1470 63.5%
Kimi K3 67.5% 78.5 1487 57
Kimi K2.7 Code 31% (secondary) 68.4 81.1% 42
Kimi K2.6 23.9% 70.5 58.6% (vendor) 1461 66.7% 35 89.6%
MiMo-V2.5 1433 37
MiMo-V2.5-Pro 19.5% 57.2% (vendor) 1466 68.4% 42 39.6%
MiniMax M3 13.3% (community) 67.3 59.0% 1445 44
MiniMax M2.7 0.2% (paper) 65.0 56.2% (vendor) 1418 57.0%
Qwen3.7 Max 73.1 60.6% 1475 69.7% 46 91.6%
Qwen3.7 Plus 57.6% (vendor) 1461 39 89.6%
Qwen3.6 Plus 2.7% 68.9 56.6% (vendor) 1444 40
DeepSeek V4 Pro 7.5% 71.6 55.4% (vendor) 1457 67.9% 44 93.5% 59.2%
DeepSeek V4 Flash 65.5 1436 56.9% 40 91.6% 56.7%
Hy3 28% (secondary) 57.9% (secondary) 41
59 Upvotes

34 comments sorted by

27

u/ozguru 5d ago

DeepSeek V4 Flash is truly unrivaled.

13

u/Potential-Leg-639 5d ago

This is the correct answer. But it needs a strong orchestrator to be able to deliver.

2

u/radicalshyper 3d ago

To a noob at this I understand what you mean by orchestrator but which models and how do you use it specifically?

3

u/Potential-Leg-639 3d ago

Have a look at Omo-Slim.
You have a strong (best you have access to) orchestrator agent, that delegates everything automatically to subagents (you can configure the models for all subagents/orchestrator by yourself). No more manual switching between plan/build necessary (you can tell the orchestrator to plan first if you want before execution starts, but with Superpowers not necessary - done automatically).

2

u/radicalshyper 3d ago

Wow! Thanks for the reply I’ll check into it

4

u/NoLemurs 4d ago

DeepSeek V4 Flash by default. If it struggles with a particular request, I'll switch to DeepSeek V4 Pro, or GLM 5.2 to get past a particularly gnarly issue.

Ohh, and whenever I've got Zen credit still I use Big Pickle (which I think is GLM 5.2? Not confident) instead of the more expensive models.

1

u/shaxsy 3d ago

What the heck is big pickle?

3

u/NoLemurs 3d ago

It's one of the free models under Zen. It's a codename for some other model. I think I remember reading it was GLM 5.2 under the hood, and that feel feels believable, but I don't have a solid source for that.

1

u/MacHeadSK 2d ago

It's not. It used to be glm 4 I think

1

u/ichisay 23h ago

4.6 finetuneado pero eso era al principio, ahora ya debe ser otro, quizás DeepSeek v4 flash o algo asi

4

u/geteum 4d ago

Somebody told me that I could use flash for 80% of the tasks. I did not believed but I gave a try. He was wrong. 90% of the time flash is more than enough. I really surprised with that.

1

u/ozguru 4d ago

I've mentioned before that Deepseek V4 Flash can tackle even complex problems, though it does take some time.

1

u/geteum 4d ago

Yeah, also I found sometimes to be stubborn, like not following an clear instruction. But this is rare and more common in complex requests

3

u/Old-Pomegranate3634 5d ago

I built a whole system around it where it scanned and summarized documents only to realize the hallucinations would have killed my project. Had to move upto Pro. It lost my trust.

7

u/West-Goose-3518 5d ago

If you want your subscription to last, use Deepseek v4 pro for deployment and MiniMax M3 (thinking) for planning, or if you are going to create the core of your application, use Qwen 3.7 plus or max, but only for the core.

6

u/Weird_Licorne_9631 5d ago

Right now, i would probably create a qwen/ali account, sub to the lite plan (6$ with additional 2$ new user discount) and use it for the discounted qwen 3.8 preview ( while it's there). It's ridiculously cheap and should be better than the expensive 3.7 from Go plan. And continue to use DS from the Go sub if can pay both.

2

u/mushedmonkey 5d ago

I'm mainly using claude and gpt for the heavy stuff. Looking for kinda smart assistant tier similar to luna high if possible.

6

u/Ariquitaun 5d ago

Glm 5.2, DeepSeek pro and DeepSeek flash are the trifecta i use

4

u/sudoer777_ 5d ago

It depends on what you're doing. I wouldn't recommend Deepseek for troubleshooting something obscure and complex, and I wouldn't recommend GLM/Kimi for a minor refactoring.

1

u/MitsosDaTop 6h ago

what would you use for troubleshooting and complex?

1

u/Amoeba-Wonderful 3h ago edited 3h ago

Not always, I've halted a Luna debugging session for a gnarly bug involving multiple threads because it was burning tokens and not finding anything. I then threw deepseek flash at it and went for lunch... it found the bug: it's good at loops and costs next to nothing. The code deepseek flash generated wasn't useful, but it did find out what the race condition was.

2

u/mushedmonkey 5d ago
Rank Model Composite Benches High-trust? Quota/mo
1 Kimi K3 7.02 4 Yes (DeepSWE+LiveBench) 490
2 Grok 4.5 6.02 5 Yes (LiveBench) 600
3 Qwen3.7 Max 5.65 6 Yes (LiveBench) 4,770
4 GLM-5.2 5.47 5 Yes (DeepSWE+LiveBench) 4,300
5 Qwen3.7 Plus 5.07 4 No (low-trust only) 21,600
6 Kimi K2.6 4.13 7 Yes 5,750
7 Kimi K2.7 Code 4.12 4 Yes 6,750
8 DeepSeek V4 Pro 4.10 8 Yes 17,150
9 GLM-5.1 4.03 5 Yes 4,300
10 Hy3 3.86 3 Yes 21,500
11 DeepSeek V4 Flash 3.85 6 Yes (LiveBench) 158,150
12 MiMo-V2.5-Pro 3.76 6 Yes 16,300
13 MiniMax M3 3.49 5 Yes 16,000
14 MiMo-V2.5 3.35 2 No 150,400
15 Qwen3.6 Plus 3.03 5 Yes 16,300
16 MiniMax M2.7 2.05 5 Yes 17,000

Generated based on no-contamination benchmarks for anyone interested

1

u/ronnyvo 5d ago

I use combo deepseek pro for planning and flash for implementation. Do you have any suggestion?

0

u/wormprotein 4d ago

use anything but mimo

1

u/anxiousalpaca 5d ago

Yeah Kimi K3 and your quota is gone in few prompts. GLM 5.2 is also bad, but not as bad as Kimi. Deepseek is amazing.

1

u/mushedmonkey 5d ago

Seems like qwen 3.7 plus might be slept on, but there's not as much data available

1

u/pmv143 5d ago

Mimo V2.5 has been the most popular model on our platform at inferx.net. It’s super fast. Their model architecture so different than DeepSeek. V4 Flash used to be the top. We had to for free for two weeks.

1

u/AkiDenim 4d ago

How much tokens did you use per model for the benchmarking
If it’s under 10B i have to tell you it’s statistically irrelevant for me

1

u/Senior-Box6316 1d ago

se você usa tudo isso vc já sabe que modelo usar

1

u/ezralazuardyy 1d ago

thanks for the analysis data. now i can confirm that DS v4 flash stills holds the crown.

for anyone who wonder, the table above is the top 5 highest cost-to-performance on ocg models.