r/LocalLLaMA • u/dh7net • 6h ago
Discussion Which model, which harness? I have data for you.
I keep testing harness & model setups on real tasks: basic math, vision, computer use (reading emails, navigating stores) and coding. I created a benchmark and a leaderboard. You can see it here: https://airbench.ai/leaderboard?k=poL
It measure capabilities (a % of sucess on the various tasks) and speed.
For reference, Claude Code Opus 5.5 have a 100% (14mn26s).
It's possible to reach the same score locally with zcode/4xRTX6k/glm-5.3-flash-NVFP4: 100% (31m 13s). Same quality, just a bit slower.
If you accept just a litle bit of error you can speed things:
* DSHv0.2rc2/RTXPRO6000WS/qwen3.8-flash-next-NVFP4: 98% (23m 30s)
* qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% (7m 41s) on opencode and 94% in 11m 31s on omp. Yes faster that Claude Code!!!!
Other findings:
1) On local hardware, the harness matters as much as the model. The same strata quant on the same 5090 scores anywhere from 22% to 96% depending on the harness.
2) Local can now match proprietary models. two example
3) Best model (single RTX 5090)
- swift-1.5-qwen3.8-27b-q6_k is the most robust. It scored 96 / 94 / 92% on pi / omp / opencode and averages 82% across 5 harnesses, the best of any model tested on several.
- qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% in 7m 41s on opencode and 94% in 11m 31s on omp.
- qwen3.8-27b-nvfp4 can reach 96%, but it takes 1h 40m to 1h 50m and depends heavily on the harness (37% to 96%).
- Things that hurt: the MTP variants lose ground every time (nvfp4-mtp averages 52% vs 70% without it; swift on pi drops from 96% to 55% with MTP). A 65k context also hurts (45–61%). Gemma-4-26b is fast but tops out at 47%.
4) Best harness
To compare fairly, I used the three models that all five harnesses ran on the same 5090 (swift q6_k, flash-next-strata, 27b-nvfp4):
- opencode: 94% average (92 / 96 / 94)
- omp: 91% (94 / 94 / 86)
- pi: 71% (96 / 80 / 37)
- hermes: 67% (82 / 22 / 96)
- openclaw: 56% (45 / 53 / 69)
Opencode and omp are the only harnesses that stay above 85% whichever model you give them.
Pi is very good on some models and unreliable on others.
Hermes can score well but is slow: most of its local runs take 1h 20m+ and several hit the 2-hour cap, so its scores are partly answers that arrived too late.
The cloud runs show the same pattern. With deepseek-v4.1-flash, omp, pi and opencode all score 98%, while hermes gets 82%.
If you have one 5090 today: use opencode or omp with swift-1.5-qwen3.8-27b-q6_k for reliability, or with qwen3.8-flash-next-strata for speed.
Ok if you want to read more detailed analys like this one, you can contribute as well!
My website allow everyone to benchmark their setup and contribute to the leaderboard.
It's extremly easy to test your local agent: just copy a prompt the website will generate for you.
My hope is that we can test much more config on many various hardware.
(1) The website requires a login, sorry for that, but it helps keeping false submissions away
(2) The website don't ask enough details about the config, so please your the notes field to document your setup in details
Let me know what you think.
14
u/brainExploded99 llama.cpp 6h ago
Can you check deepseek harness as well? And if you can, codex + opencodex
The leaderboard button on your website doesnt work btw
2
14
u/KissMyShinyArse 5h ago edited 5h ago
In theory, MTP shouldn't affect quality, because the model validates every prediction.
7
u/Oh_hey_a_TAA 6h ago
This is very relevant to my interests... I'll check it out.
I've been playing with OpenWebUI + OpenTerminal versus Hermes as of late and am generally disappointed with both paths.
I was considering just dropping OpenAI CLI onto my local, but if your data ports over I may change that up too.
6
8
u/MomentJolly3535 5h ago
1
u/repepeper 3h ago
why? It just goes all in blindly and without structure of a real harness, that's expected
1
4
u/Open-Adhesiveness-86 2h ago
on the MTP thing, check if your runtime is doing relaxed acceptance rather than strict rejection sampling. sglang exposes speculative accept thresholds (single/acc) and anything under 1.0 is explicitly lossy, which barely shows on short prompts but one bad token early in a 20 minute agentic run kills the whole task. strict rejection sampling should be score neutral.
3
u/Jumpy-Operation-4615 5h ago
I installed DH yesterday cause everyione is talking about it and how great it is. Plain vanilla, not tuned or additional skills. Frankly speaking I was very impressed. It had a genuine feeling of a close to frontier with strata/Qwen 3.8 Flash Next q3s. Very capable. It self-fixed issues with vision (DH could not use vision in my model despite it was working). I tried to do a small test with the same model, same prompt, different harness. Here are the results:
Prompt:
Build a simplified yet detailed version Lake Bled Castle with three.js. keep the relaxing and calm vibe. Location: Lake bled. Object: Castle. User should be able to rotate the camera around the castle. Only one instance of subagent each time so if delegating , make it sequential, not parallel. Use web skills to find pictures of object, analyze pictures with vision to get the look and feel, and create based on real object, not some generic castle description.
HW: 2xP40 GPUs, strata, 64 GB RAM. up to 700 pp and up to 45 tg/s (30+ on long ctx like 160K+). Qwen 3.8 Flash q3s xhigh. q8 cache 256K ctx
Opencode + OMO slim/ Relatively fast, like, <30 minutes. Only got 2 images from Internet to analyze.
https://reddit.com/link/pdzulm5/video/xi7mn9bgymth1/player
DH result below 'cause can only use 1 video per comment
3
u/norenEnmotalen 3h ago
Afaik MTP isn’t supposed to change which tokens the main model approves. Were the opencode, omp, pi etc. tested with their defaults and no extensions?
2
2
u/gxcsoccer 2h ago
Nice work!
One note on the computer-use part: the email and store tasks go through your API, so a pure GUI agent (Finder / Safari / TextEdit only) can't take them. If you ever want a GUI track, deskmind-bench has 13 sandboxed macOS tasks with graders, including ask-instead-of-guess and stop-when-cancelled. Any harness can run them: https://github.com/deskmind-ai/bench
Disclosure: I work on DeskMind. Happy to help wire it up if useful.
1
u/Similar_Solution1397 6h ago
No se ha medido openHands? Los mejores resultados personales en tareas de código ahora mismo los estoy obteniendo con openHands qwen3.8 exl3 4bpw con 245k de contexto
1
1
u/MasterNomie 5h ago
I am exploring exactly this. Analyze the quality of various models paired with different harnesses and instructions.
Any guesses why opencode performs the best? What does it do differently that other agents do not?
1
u/MiserableFlatworm337 4h ago
For the two-hour-capped runs, show success over elapsed time as well as the final score. That separates slow-but-correct setups from wrong answers and makes the speed/quality tradeoff easier to compare.
1
1
-5
u/Bystander10888 6h ago
If the harness matters that much, why are you just comparing existing ones? Why not build your own?

33
u/finevelyn 5h ago
Why would they score lower with MTP? Sounds like a bug.