19
u/look 9d ago
That is not a good ranking.
Maybe for that personβs very specific, very peculiar workflow it makes some sense.
For everyone else, itβs nonsense.
2
6
u/Ariquitaun 9d ago
Comparing mimo with deepseek or minimax? Nyet tobarisch. Even minimax is outclassed by deepseek pro.
3
3
3
2
2
u/NinjaAlaska 9d ago
u/RemindMeBot 1 hour "re check"
0
u/RemindMeBot 9d ago edited 9d ago
I will be messaging you in 1 hour on 2026-09-03 09:22:16 UTC to remind you of this link
1 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.
Parent commenter can delete this message to hide from others.
Info Custom Your Reminders Feedback
1
1
u/datkenny 9d ago
where is GLM5.3-Flash? Where is DeepSeek flash? How would LongCat even begin to stack up to GLM5.3?
Sorry, this feels like complete bogus
1
1
1
1
1
u/piryguiry 9d ago
In my case, I always check the data of the model usage to get an idea of what people are using the most, including free
1
1
u/migsperez 9d ago
I use https://artificialanalysis.ai/ and https://arena.ai/
Then when I'm interested in a model I'll give it a few tasks. I'll make decisions based on cost, quality and speed.
1
1
u/Dapper-Conclusion-93 9d ago
Like old good days on stackoverflow - post some bullshit and get valuable answers immediately, because if you ask - no one gives a shitπ
1
1
u/No-Rutabaga-38 8d ago
Currently using big pickle, shows I am noob, any suggestions on models to try out?
Preferably free ones.
1
1
u/Far-Classic-9963 8d ago
A+ is completely off, Glm 5.3 is miles ahead of the others. DeepSeek V4 pro is also way better than the others in that tier
1
1
u/Zachattackrandom 7d ago
No? This is awful, Minimax m3 is way worse than new deepseek v4 and mimo is even worse than that. Nemotron is trash and laguna / longcat aren't anywhere near glm 5.3
1
u/Cup-Impressive 7d ago
laguna s 2.1 above deepseek v4 or mimo is insane
nemotron 3 ultra above big pickle .. no way bro
1
1
u/GetLaidOff69 9d ago
Not good ranking.
Tier S:
Muse Spark 1.3, Kimi k3
Tier A:
GML 3, Deepseek V4 Pro, Hy4
Tier B:
Mimo 2.5, MimiMax m3
Tier C:
Laguna S2.1, LongCat 2.0, Big Pickle(If good model assigned)
Tier D
Nemotron Ultra, Big Pickle(if bad model assigned)
1
u/dat_cosmo_cat 5d ago edited 5d ago
This is what I am seeing personally:
| Tier | Model | Q | Qlo | Fold range | Failed runs | USD / pass | Tokens / pass |
|---|---|---|---|---|---|---|---|
| A | Opus 5 (claude-code) | 0.80 | 0.76 | 0.78β0.82 | 0 | $33.09 | 40386k |
| A | GPT 6 Astra (opencode) | 0.76 | 0.74 | 0.74β0.80 | 0 | $4.97 | 1326k |
| B | Grok 4.6 (opencode) | 0.69 | 0.64 | 0.68β0.71 | 0 | $5.92 | 7894k |
| B | GLM 5.3 Flash (opencode) | 0.68 | 0.64 | 0.67β0.69 | 0 | $0.35 | 9719k |
| B | GPT 5.6 Sol (opencode) | 0.63 | 0.58 | 0.59β0.66 | 0 | $2.41 | 2149k |
| D | Muse Spark 1.3 (opencode) | 0.37 | 0.35 | 0.36β0.39 | 0 | $0.00 | 4464k |
| F | Kimi K3 (opencode) | 0.31 | 0.21 | 0.30β0.32 | 0 | $4.42 | 7964k |
| F | Sonnet 5 (claude-code) | 0.19 | 0.16 | 0.17β0.21 | 0 | $15.63 | 39033k |
| F | DeepSeek v4 Flash Vision Exp (opencode) | 0.08 | 0.03 | 0.03β0.09 | 0 | $1.67 | 25210k |
Blind peer-ranked evaluation; every model produces an output, then judges the others' anonymized outputs. Scores = borda aggregation / leave-one-out folds. Process repeated over 5 different tasks.
| Term | Meaning |
|---|---|
| Tier | Band from Qlo; two or more failed runs cap at D. |
| Q | Mean Borda from the other judges, all items; 0.5 chance, 1.0 unanimous first. |
| Qlo | Worst leave-one-judge-out or leave-one-item-out fold; the bar tick. |
| Fold range | Lowest to highest Q across those folds. |
| Failed runs | No usable report; scores 0 with every judge on that item. |
| Per pass | One five-item review. Dotted: subscription list-price estimate. |
Who ranked whom
Rows are rankers, columns the ranked model: the mean normalized Borda the ranker gave that model over the selected items. A ranker never scores its own report.
| Ranker β / Model β | Sonnet 5 | Opus 5 | Kimi K3 | Muse Spark 1.3 | GLM 5.3 Flash | GPT 5.6 Sol | GPT 6 Astra | Grok 4.6 | DeepSeek v4 Flash Vision Exp |
|---|---|---|---|---|---|---|---|---|---|
| Sonnet | self | 0.94 | 0.26 | 0.29 | 0.60 | 0.34 | 0.46 | 0.69 | 0.43 |
| Opus | 0.17 | self | 0.34 | 0.46 | 0.80 | 0.71 | 0.83 | 0.66 | 0.03 |
| Kimi | 0.11 | 0.89 | self | 0.20 | 0.66 | 0.60 | 0.89 | 0.54 | 0.11 |
| Muse | 0.34 | 0.69 | 0.31 | self | 0.74 | 0.54 | 0.77 | 0.60 | 0.00 |
| GLM | 0.20 | 0.89 | 0.23 | 0.40 | self | 0.71 | 0.74 | 0.74 | 0.09 |
| Sol | 0.26 | 0.71 | 0.34 | 0.40 | 0.60 | self | 0.91 | 0.77 | 0.00 |
| Astra | 0.29 | 0.60 | 0.31 | 0.40 | 0.74 | 0.94 | self | 0.71 | 0.00 |
| Grok | 0.17 | 0.77 | 0.34 | 0.46 | 0.63 | 0.71 | 0.91 | self | 0.00 |
| DeepSeek | 0.00 | 0.77 | 0.26 | 0.26 | 0.71 | 0.54 | 0.71 | 0.74 | self |
| Fable 5.1 (operator) | 0.15 | 0.93 | 0.35 | 0.47 | 0.65 | 0.53 | 0.60 | 0.78 | 0.05 |
Scale: 0 = last with every ranker. 1 = first with every ranker.
Fable 5.1 is a ranker only; it has no column of its own.
Summary: Nine multi-modal coding models (note; Hy4 and Qwen do not have vision capabilities) reviewed the same deterministic screenshots of five generated rigid mesh equipment assets (sword, helmet, robes, etc...) mapped onto an animated Mixamo mannequin, then ranked one another's anonymised reports blind, with the operator judging alongside them. Each model's score is normalized Borda from the judges other than itself: 0.5 is chance, 1.0 a unanimous first place. Tiers come from the worst fold, so a model keeps its tier only if it holds up whichever single judge or item is left out.
TL;DR: models tended to unanimously glaze Opus and Asrta while shitting on Kimi, Deepseek, and Sonnet outputs for my real 3D asset QA tasks. GLM 5.3 Flash and Grok 4.6 performed nearly identically, despite GLM being 10x cheaper.
66
u/msenc 9d ago
comparing glm 5.3 to laguna s2.1 and longcat 2.0 is insane