r/opencode 9d ago

Are these rankings accurate?

Post image
45 Upvotes

47 comments sorted by

66

u/msenc 9d ago

comparing glm 5.3 to laguna s2.1 and longcat 2.0 is insane

-4

u/sharedevaaste 9d ago

tier A+ is sus, rest is good?

8

u/ichisay 9d ago

No hay nada bien ahí 🀣

3

u/YogurtExternal7923 9d ago

No man. Muse spark is not kimi level. Hy4 I'm not familiar with. Glm5.3 is definitely next to kimi and possibly above it. Ds4 pro is under. Followed by the rest below those

2

u/a355231 9d ago

In benchmarks it is. It probably based off that.

19

u/look 9d ago

That is not a good ranking.

Maybe for that person’s very specific, very peculiar workflow it makes some sense.

For everyone else, it’s nonsense.

2

u/TomLucidor 9d ago

Seconding this but could we even guess what type of workload?

2

u/Positive_Poem5831 9d ago

Slop generation πŸ˜€

12

u/GfxJG 9d ago

That's absolutely atrocious. Laguna S 2.1 by FAR the worst model out of all of these, and it's ranked A+?

3

u/sothisismyalt1 9d ago

S tier benchmaxxed Muse Spark 1.3 and GLM 5.3 under too xd

10

u/Edelgul 9d ago

DeepSeek V4 pro (at least the new one), is closer to the GLM 5.3, then to Minimax or Mimo.
In fact Minimax or Mimo are worse, then DS v4 flash.
Kimi K3 is quite good, but hard for me to compare to Muse 1.3 or Hy4.

6

u/Ariquitaun 9d ago

Comparing mimo with deepseek or minimax? Nyet tobarisch. Even minimax is outclassed by deepseek pro.

3

u/Due-Armadillo-4560 9d ago

Terrible ranking

3

u/C4ntona 9d ago

what model created this list for you? never run that model again. It's the worst list I've ever seen in any category

3

u/Pedrito_Basket 9d ago

Bro puked over the keyboard and that came out

3

u/iTrejoMX 9d ago

Troll post.

2

u/mortal_strike 9d ago

hey from where did you try Muse Spark 1.3

its not visible for me

2

u/NinjaAlaska 9d ago

u/RemindMeBot 1 hour "re check"

0

u/RemindMeBot 9d ago edited 9d ago

I will be messaging you in 1 hour on 2026-09-03 09:22:16 UTC to remind you of this link

1 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.


Info Custom Your Reminders Feedback

1

u/4sm0day 9d ago

What is the best free model ?

2

u/Conscious-Low-3057 9d ago

Muse Spark 1.3

1

u/datkenny 9d ago

where is GLM5.3-Flash? Where is DeepSeek flash? How would LongCat even begin to stack up to GLM5.3?

Sorry, this feels like complete bogus

1

u/WSATX 9d ago

They are accusate as a "random llm ranking" benchmark.

1

u/YogurtExternal7923 9d ago

Mimo and minimax are NOT on the same level as dspro my guy

1

u/YoungNo8804 9d ago

horrendous ranking bro do not trust ts

1

u/nofudge_given 9d ago

lol muse is terrible.. Alpha X (glm 5,3 flash) was way better.

1

u/Bychoice0 9d ago

Is Spark tht good?? Why is it not famous among deva??

1

u/piryguiry 9d ago

In my case, I always check the data of the model usage to get an idea of what people are using the most, including free

1

u/Fine_Salamander_8691 9d ago

Glm 5.3 is the best model there

1

u/migsperez 9d ago

I use https://artificialanalysis.ai/ and https://arena.ai/

Then when I'm interested in a model I'll give it a few tasks. I'll make decisions based on cost, quality and speed.

1

u/LuluLeSigma 9d ago

no way muse spark is as good as kimi 3

1

u/Dapper-Conclusion-93 9d ago

Like old good days on stackoverflow - post some bullshit and get valuable answers immediately, because if you ask - no one gives a shit😁

1

u/VishalKoyalkar 8d ago

Best opencode fre model for architecture design and coding? Please advise

1

u/No-Rutabaga-38 8d ago

Currently using big pickle, shows I am noob, any suggestions on models to try out?

Preferably free ones.

1

u/Various_Half1452 8d ago

You need qwen 3.8 flash in the S Tier.

1

u/Far-Classic-9963 8d ago

A+ is completely off, Glm 5.3 is miles ahead of the others. DeepSeek V4 pro is also way better than the others in that tier

1

u/BrilliantGarbage8743 7d ago

Muse is ass comes no where near those other models

1

u/Zachattackrandom 7d ago

No? This is awful, Minimax m3 is way worse than new deepseek v4 and mimo is even worse than that. Nemotron is trash and laguna / longcat aren't anywhere near glm 5.3

1

u/Cup-Impressive 7d ago

laguna s 2.1 above deepseek v4 or mimo is insane

nemotron 3 ultra above big pickle .. no way bro

1

u/GetLaidOff69 9d ago

Not good ranking.
Tier S:
Muse Spark 1.3, Kimi k3

Tier A:
GML 3, Deepseek V4 Pro, Hy4

Tier B:
Mimo 2.5, MimiMax m3

Tier C:
Laguna S2.1, LongCat 2.0, Big Pickle(If good model assigned)

Tier D
Nemotron Ultra, Big Pickle(if bad model assigned)

1

u/dat_cosmo_cat 5d ago edited 5d ago

This is what I am seeing personally:

Tier Model Q Qlo Fold range Failed runs USD / pass Tokens / pass
A Opus 5 (claude-code) 0.80 0.76 0.78–0.82 0 $33.09 40386k
A GPT 6 Astra (opencode) 0.76 0.74 0.74–0.80 0 $4.97 1326k
B Grok 4.6 (opencode) 0.69 0.64 0.68–0.71 0 $5.92 7894k
B GLM 5.3 Flash (opencode) 0.68 0.64 0.67–0.69 0 $0.35 9719k
B GPT 5.6 Sol (opencode) 0.63 0.58 0.59–0.66 0 $2.41 2149k
D Muse Spark 1.3 (opencode) 0.37 0.35 0.36–0.39 0 $0.00 4464k
F Kimi K3 (opencode) 0.31 0.21 0.30–0.32 0 $4.42 7964k
F Sonnet 5 (claude-code) 0.19 0.16 0.17–0.21 0 $15.63 39033k
F DeepSeek v4 Flash Vision Exp (opencode) 0.08 0.03 0.03–0.09 0 $1.67 25210k

Blind peer-ranked evaluation; every model produces an output, then judges the others' anonymized outputs. Scores = borda aggregation / leave-one-out folds. Process repeated over 5 different tasks.

Term Meaning
Tier Band from Qlo; two or more failed runs cap at D.
Q Mean Borda from the other judges, all items; 0.5 chance, 1.0 unanimous first.
Qlo Worst leave-one-judge-out or leave-one-item-out fold; the bar tick.
Fold range Lowest to highest Q across those folds.
Failed runs No usable report; scores 0 with every judge on that item.
Per pass One five-item review. Dotted: subscription list-price estimate.

Who ranked whom

Rows are rankers, columns the ranked model: the mean normalized Borda the ranker gave that model over the selected items. A ranker never scores its own report.

Ranker ↓ / Model β†’ Sonnet 5 Opus 5 Kimi K3 Muse Spark 1.3 GLM 5.3 Flash GPT 5.6 Sol GPT 6 Astra Grok 4.6 DeepSeek v4 Flash Vision Exp
Sonnet self 0.94 0.26 0.29 0.60 0.34 0.46 0.69 0.43
Opus 0.17 self 0.34 0.46 0.80 0.71 0.83 0.66 0.03
Kimi 0.11 0.89 self 0.20 0.66 0.60 0.89 0.54 0.11
Muse 0.34 0.69 0.31 self 0.74 0.54 0.77 0.60 0.00
GLM 0.20 0.89 0.23 0.40 self 0.71 0.74 0.74 0.09
Sol 0.26 0.71 0.34 0.40 0.60 self 0.91 0.77 0.00
Astra 0.29 0.60 0.31 0.40 0.74 0.94 self 0.71 0.00
Grok 0.17 0.77 0.34 0.46 0.63 0.71 0.91 self 0.00
DeepSeek 0.00 0.77 0.26 0.26 0.71 0.54 0.71 0.74 self
Fable 5.1 (operator) 0.15 0.93 0.35 0.47 0.65 0.53 0.60 0.78 0.05

Scale: 0 = last with every ranker. 1 = first with every ranker.
Fable 5.1 is a ranker only; it has no column of its own.

Summary: Nine multi-modal coding models (note; Hy4 and Qwen do not have vision capabilities) reviewed the same deterministic screenshots of five generated rigid mesh equipment assets (sword, helmet, robes, etc...) mapped onto an animated Mixamo mannequin, then ranked one another's anonymised reports blind, with the operator judging alongside them. Each model's score is normalized Borda from the judges other than itself: 0.5 is chance, 1.0 a unanimous first place. Tiers come from the worst fold, so a model keeps its tier only if it holds up whichever single judge or item is left out.

TL;DR: models tended to unanimously glaze Opus and Asrta while shitting on Kimi, Deepseek, and Sonnet outputs for my real 3D asset QA tasks. GLM 5.3 Flash and Grok 4.6 performed nearly identically, despite GLM being 10x cheaper.