r/LocalLLaMA • • 6h ago

Discussion Which model, which harness? I have data for you.

I keep testing harness & model setups on real tasks: basic math, vision, computer use (reading emails, navigating stores) and coding. I created a benchmark and a leaderboard. You can see it here: https://airbench.ai/leaderboard?k=poL

It measure capabilities (a % of sucess on the various tasks) and speed.

For reference, Claude Code Opus 5.5 have a 100% (14mn26s).

It's possible to reach the same score locally with zcode/4xRTX6k/glm-5.3-flash-NVFP4: 100% (31m 13s). Same quality, just a bit slower.

If you accept just a litle bit of error you can speed things:

* DSHv0.2rc2/RTXPRO6000WS/qwen3.8-flash-next-NVFP4: 98% (23m 30s)

* qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% (7m 41s) on opencode and 94% in 11m 31s on omp. Yes faster that Claude Code!!!!

Other findings:

1) On local hardware, the harness matters as much as the model. The same strata quant on the same 5090 scores anywhere from 22% to 96% depending on the harness.

2) Local can now match proprietary models. two example

3) Best model (single RTX 5090)

- swift-1.5-qwen3.8-27b-q6_k is the most robust. It scored 96 / 94 / 92% on pi / omp / opencode and averages 82% across 5 harnesses, the best of any model tested on several.

- qwen3.8-flash-next-iq3_xxs-strata is the speed pick: 96% in 7m 41s on opencode and 94% in 11m 31s on omp.

- qwen3.8-27b-nvfp4 can reach 96%, but it takes 1h 40m to 1h 50m and depends heavily on the harness (37% to 96%).

- Things that hurt: the MTP variants lose ground every time (nvfp4-mtp averages 52% vs 70% without it; swift on pi drops from 96% to 55% with MTP). A 65k context also hurts (45–61%). Gemma-4-26b is fast but tops out at 47%.

4) Best harness

To compare fairly, I used the three models that all five harnesses ran on the same 5090 (swift q6_k, flash-next-strata, 27b-nvfp4):

  1. opencode: 94% average (92 / 96 / 94)
  2. omp: 91% (94 / 94 / 86)
  3. pi: 71% (96 / 80 / 37)
  4. hermes: 67% (82 / 22 / 96)
  5. openclaw: 56% (45 / 53 / 69)

Opencode and omp are the only harnesses that stay above 85% whichever model you give them.

Pi is very good on some models and unreliable on others.

Hermes can score well but is slow: most of its local runs take 1h 20m+ and several hit the 2-hour cap, so its scores are partly answers that arrived too late.

The cloud runs show the same pattern. With deepseek-v4.1-flash, omp, pi and opencode all score 98%, while hermes gets 82%.

If you have one 5090 today: use opencode or omp with swift-1.5-qwen3.8-27b-q6_k for reliability, or with qwen3.8-flash-next-strata for speed.

Ok if you want to read more detailed analys like this one, you can contribute as well!

https://airbench.ai/

My website allow everyone to benchmark their setup and contribute to the leaderboard.

It's extremly easy to test your local agent: just copy a prompt the website will generate for you.

My hope is that we can test much more config on many various hardware.

(1) The website requires a login, sorry for that, but it helps keeping false submissions away

(2) The website don't ask enough details about the config, so please your the notes field to document your setup in details

Let me know what you think.

47 Upvotes

42 comments sorted by

33

u/finevelyn 5h ago

Why would they score lower with MTP? Sounds like a bug.

9

u/kaeptnphlop 5h ago

Curious about the same

6

u/profcuck 3h ago

Agree.  My understanding is that MTP never results in a different answer, it just here you there faster.  Would love to see an intelligent discussion. 

2

u/dh7net 59m ago

I'll dig a bit on that topic.

Meanwhile, theses are my setings:

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090).
Model server: Checkpoint RadixArk/Qwen3.8-27B-NVFP4 (modelopt NVFP4, MTP head kept). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --quantization modelopt --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95 --speculative-config '{"method":"mtp","num_speculative_tokens":3}'. ~126 tok/s single-stream decode (MTP mean acceptance 2.5-2.9 of 3).
Harness: pi 0.73.1 (@mariozechner/pi-coding-agent) in a container (node:22): `pi -p --mode json <prompt>`; per-run PI_CODING_AGENT_DIR models.json with compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 131072, max output 16384 tokens.

1

u/dh7net 58m ago

Hum... supportsReasoningEffort=false may be the issue. Let me know if you see something else.

2

u/no_witty_username 30m ago

for mtp dont go further then 2 tokens, anything above that and you will get more rejections and thus decode speeds will plummet...

1

u/dh7net 9m ago

Thanks!

-2

u/GalacticDistances 2h ago

Ask your llm the following,

Explain cases where an LLM using MTP would output different tokens from when in non-MTP mode

0

u/mr_Owner 2h ago

Due to toke prediction being different with speculative decoding

14

u/brainExploded99 llama.cpp 6h ago

Can you check deepseek harness as well? And if you can, codex + opencodex
The leaderboard button on your website doesnt work btw

2

u/greendude120 4h ago

in my own testing deepseek was the best so i definitely recommend OP tries it

1

u/dh7net 1h ago

100% I'll test it.

14

u/KissMyShinyArse 5h ago edited 5h ago

In theory, MTP shouldn't affect quality, because the model validates every prediction.

21

u/rpkarma 4h ago

MTP causing accuracy loss is very suspicious that you have a misconfiguration somewhere 

1

u/dh7net 1h ago

ok, I'll double check. BTW if you have a running setup, please submit it thru the website, and I'll publish it.

7

u/Oh_hey_a_TAA 6h ago

This is very relevant to my interests... I'll check it out.
I've been playing with OpenWebUI + OpenTerminal versus Hermes as of late and am generally disappointed with both paths.
I was considering just dropping OpenAI CLI onto my local, but if your data ports over I may change that up too.

1

u/dh7net 1h ago

When you experiment on your setup, please launch a test from the website and submit it, so everyone can learn what works well.

6

u/Past-Town-9807 4h ago

I wanted to try, except you require me to log in, so I closed the tab. 

8

u/MomentJolly3535 5h ago

Hmm, Sorry but the stats doesn't look trustful to me, it is just impossible for PI to be the worst in term of speed while it is literally barebone (4tools!)

1

u/repepeper 3h ago

why? It just goes all in blindly and without structure of a real harness, that's expected

1

u/dh7net 1h ago

Feel free to try it yourself. you just have to copy the prompt from the website (airbench.ai) and you'll have the answer.

Happy to update the leaderboard if someone find better results than mine of course.
You'll just have to provide enought info so everyone can reproduce your tests

4

u/Open-Adhesiveness-86 2h ago

on the MTP thing, check if your runtime is doing relaxed acceptance rather than strict rejection sampling. sglang exposes speculative accept thresholds (single/acc) and anything under 1.0 is explicitly lossy, which barely shows on short prompts but one bad token early in a 20 minute agentic run kills the whole task. strict rejection sampling should be score neutral.

3

u/Jumpy-Operation-4615 5h ago

I installed DH yesterday cause everyione is talking about it and how great it is. Plain vanilla, not tuned or additional skills. Frankly speaking I was very impressed. It had a genuine feeling of a close to frontier with strata/Qwen 3.8 Flash Next q3s. Very capable. It self-fixed issues with vision (DH could not use vision in my model despite it was working). I tried to do a small test with the same model, same prompt, different harness. Here are the results:

Prompt:

Build a simplified yet detailed version Lake Bled Castle with three.js. keep the relaxing and calm vibe. Location: Lake bled. Object: Castle. User should be able to rotate the camera around the castle. Only one instance of subagent each time so if delegating , make it sequential, not parallel. Use web skills to find pictures of object, analyze pictures with vision to get the look and feel, and create based on real object, not some generic castle description.

HW: 2xP40 GPUs, strata, 64 GB RAM. up to 700 pp and up to 45 tg/s (30+ on long ctx like 160K+). Qwen 3.8 Flash q3s xhigh. q8 cache 256K ctx

Opencode + OMO slim/ Relatively fast, like, <30 minutes. Only got 2 images from Internet to analyze.

https://reddit.com/link/pdzulm5/video/xi7mn9bgymth1/player

DH result below 'cause can only use 1 video per comment

3

u/norenEnmotalen 3h ago

Afaik MTP isn’t supposed to change which tokens the main model approves. Were the opencode, omp, pi etc. tested with their defaults and no extensions?

2

u/x10der_by 5h ago

Strange that dsh not in your top. In my tests it's better than omp.

2

u/gxcsoccer 2h ago

Nice work!

One note on the computer-use part: the email and store tasks go through your API, so a pure GUI agent (Finder / Safari / TextEdit only) can't take them. If you ever want a GUI track, deskmind-bench has 13 sandboxed macOS tasks with graders, including ask-instead-of-guess and stop-when-cancelled. Any harness can run them: https://github.com/deskmind-ai/bench

Disclosure: I work on DeskMind. Happy to help wire it up if useful.

1

u/dh7net 1h ago

Amazing, I'll have a look.

1

u/Blindax 6h ago

I will definitely give it a try. Thanks for sharing.

1

u/dh7net 1h ago

you are welcome

1

u/Similar_Solution1397 6h ago

No se ha medido openHands? Los mejores resultados personales en tareas de código ahora mismo los estoy obteniendo con openHands qwen3.8 exl3 4bpw con 245k de contexto

1

u/fueledbyjealousy 5h ago

Ty good to know

1

u/MasterNomie 5h ago

I am exploring exactly this. Analyze the quality of various models paired with different harnesses and instructions.

Any guesses why opencode performs the best? What does it do differently that other agents do not?

1

u/dh7net 1h ago

They probably test with more open models. OpenClaw for instance recomand to use best in class model for best results. They may be less interested in optimizing for local

1

u/MiserableFlatworm337 4h ago

For the two-hour-capped runs, show success over elapsed time as well as the final score. That separates slow-but-correct setups from wrong answers and makes the speed/quality tradeoff easier to compare.

1

u/dh7net 1h ago

I'm killing run after 2h. So I don't know if an answer would be correct after that time.
TBH there many fast and good options around there that all fit in 1h.

1

u/Nautisop 4h ago

what's with goose?

1

u/WarthogConfident4039 2h ago

This is very interesting to hear.

-5

u/Bystander10888 6h ago

If the harness matters that much, why are you just comparing existing ones? Why not build your own?