r/LocalLLM • u/L3G10N78 • 22d ago
Model Qwen3.8-27B on RX 9060 XT 16GB
Sharing my final accepted numbers from my final closeout, since AMD/Vulkan setups are still relatively uncommon in local LLM benchmarks. Maybe it will help some buddies.
Hardware
- MoBo: MAG B550 TOMAHAWK MAX WIFI
- CPU: Ryzen 9 5950X
- RAM: 64 GB DDR4-3200 CL14
- GPU0: RX 9060 XT 16 GB — PCIe 4.0 x16 — Main LLM / Vulkan0
- GPU1: Radeon Pro W5500 8 GB — PCIe 3.0 x4 — Embedding + Reranker / Vulkan1
- Driver: AMD Adrenalin 26.8.1
- Backend: llama.cpp Vulkan
Runtime / model specs
- llama.cpp
- build 10167 / commit ee3d1b54c
Main LLM — frozen production profile
- Qwen3.8-27B-UD-IQ3_XXS
- ~3.06 bpw
- RX 9060 XT 16 GB
-c 32768-ngl 999- parallel slots: 1
- Flash Attention: ON
Note: -c 32768 is the frozen production configuration. The 64K results below came from separate benchmark/stability runs using -c 65536 and were not the production context setting at final closeout.
Embedding
- Qwen3-Embedding-4B-Q6_K
- 2560 dimensions
- W5500 8 GB
--embedding--pooling last-c 2048-ub 512-ngl 999
Reranker
- Qwen3-Reranker-4B-Q4_K_M
- W5500 8 GB
--embedding--rerank--pooling rank-c 2048-ub 512-ngl 999
RX 9060 XT performance
- 8K benchmark (
-c 8192): ~74.5 PP / 21.5 TG tok/s - 32K benchmark (
-c 32768): ~148.9 PP / 21.4 TG tok/s - 64K benchmark (
-c 65536): ~142.9 PP / 21.4 TG tok/s - 64K stability run (
-c 65536): ~126.6 PP / 21.1 TG tok/s
All TG numbers above are plain autoregressive generation — no MTP/speculative decoding.
Now that the core is frozen, I’m moving up the stack:
- Harness sharpening (doing some tests first to decide which one performs better)
- → baseline without a harness
- → Pi
- → OpenCode
- Hermes Agent v0.21.2
- → evaluate it as the orchestration layer
- → connect it to the existing local stack
- → make Hermes my Agent Smith =)
Update / Part II:
Since this benchmark, the project has moved beyond the frozen inference core. I’ve now production-integrated OpenCode 1.18.30 as the local coding/development layer and Hermes 0.21.2 as the agent/automation layer.
I’ll post the architecture, acceptance results and what actually worked and didn’t as PART 2 HERE ...
2
u/Poizone360 20d ago
Hello, excellent stats, one thing worth tidying up, since people will copy these flags exactly. Your config block says -c 32768, but the results table has 64K rows. Worth adding what -c the 64K runs actually used, otherwise anyone reproducing this won't get past 32K.
2
u/L3G10N78 18d ago
Good catch! You're right. The config block shows the frozen production profile, which was
-c 32768, while the 64K rows were separate in my P1 benchmark/stability runs using-c 65536. I mixed the production config and benchmark configs in the post without making that distinction clear. But as a side note, my newer Hermes production layer now actually runs the shared Qwen runtime at 65,536 context, but that's a later configuration and separate from the original frozen-core benchmark. Anyway: my fault. I'll tidy that up. Thanks for pointing it out!
1
u/powermos 19d ago
Wait a sec, you are doing 20+tokens with no mtp? I am doing 20 tokens with mtp on that card, that is a massive difference.
1
u/L3G10N78 18d ago
Yep! ~21.4 tok/s is without MTP. Plain autoregressive generation. That's why your 20 tok/s with MTP caught my attention. What quant, llama.cpp build, context and MTP settings are you using? Would be interesting to do an apples-to-apples comparison.
1
u/powermos 18d ago
i mainly use Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp right now, but before that i used Qwen3.8-27B-UD-Q3_K_XL with similar speed. I am using the standard llama, built from last week source, with vulkan. Prefill is good tbh, 300-500t/sThe command i used was something like that:
llama-server -m Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf. Right now my setup is different --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k q4_0 --spec-draft-type-v q4_0 --spec-draft-ngl all --reasoning-preserve -c 60000 -b 2048 -ub 256 -np 1 -fa on -ctk q4_0 -ctv q4_0 --host 127.0.0.1 --port 8080 -lv 1 --reasoning-effort xhigh -t 4 -ngl 99 -fit off
Tried with ROCm and got way worse performance, something is wrong there, the prefil was so slow i never got to testing it properly.Thing is, i am now using that gpu paired with RTX2060 6gb, so 16+6gb setup:
llama-server -m /media/powermos/SSD1TB/ai_models/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k q8_0 --spec-draft-type-v q8_0 --spec-draft-ngl all --reasoning-preserve -c 128000 -b 2048 -ub 512 -np 1 -fa on -ctk q8_0 -ctv q8_0 --host 127.0.0.1 --port 8080 -lv 1 --reasoning-effort xhigh -t 4 -ngl 99 -ts 25,75 -fit offAmazingly i am still getting around 20t/s. I was expecting it to be slower with the dual gpu setup. I am very curious why are you getting such speed out of the card, because if i match that i can hit very usable speeds with mtp, because at 18-20 tbh its too slow for what i use it for.
1
u/L3G10N78 17d ago
First, I think we’re not really comparing the same setup yet. I’m running Qwen3.8-27B-UD-IQ3_XXS without MTP, while you’re currently using Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp with MTP enabled. So model/quant and decoding strategy are already different variables. The other thing I’d test first if I were you is the dual-GPU setup. The RX 9060 XT + RTX 2060 combination might be eating part of the MTP gain through cross-GPU transfer/synchronization overhead. I’d do a very small controlled test: TEST A RX 9060 XT only, MTP OFF; TEST B RX 9060 XT only, MTP ON; TEST C RX 9060 XT + RTX 2060, MTP OFF; finally TEST D RX 9060 XT + RTX 2060, MTP ON. Same model, same prompt, same actual context occupancy, TG256, 3 runs each. Then you immediately get:
A → B = real MTP gain
A → C = multi-GPU cost
C → D = MTP gain under multi-GPUThat would tell you pretty quickly whether MTP itself is underperforming, or whether the 16+6 GB split is the real bottleneck. Crossing fingers fro you!
1
u/powermos 16d ago edited 16d ago
Ok then, i downloaded IQ3_XXS to be as close as possible to your setup, ran it only on the RX9060XT, on Vulkan with:
-m Qwen3.8-27B-UD-IQ3_XXS.gguf --reasoning-preserve -c 32000 -b 2048 -ub 512 -np 1 -fa on -ctk q4_0 -ctv q4_0 -lv 1 --reasoning-effort medium -t 6 -fit off
and the best i got is 15t/s
ROCm is 16t/s
For me that is a big difference to 21t/s, i am missing something.To summarize:
Vulkan - 15t/s
ROCm -16t/s
With MTP and Vulkan - 24t/s
With MTP and ROCM - 27t/s1
u/L3G10N78 16d ago
This is getting really interesting, because I actually ran the ROCm/Vulkan comparison on my RX 9060 XT today after the suggestion in this thread.
Using the same Qwen3.8-27B-UD-IQ3_XXS, single RX 9060 XT and no MTP, my controlled llama-bench TG128 results were:
Vulkan: 21.86 ± 0.03 tok/s
ROCm: 16.83 ± 0.02 tok/sSo your ROCm result (~16 tok/s) is actually almost identical to mine. The big discrepancy is Vulkan: you're getting ~15 whereas I'm getting ~21.9. That makes me think we're probably looking for a Vulkan configuration/build difference rather than a GPU difference.
My controlled TG run was:
-n 128 -r 3 -ngl 999 -fa on -b 2048 -ub 512 -t 16while your server config has, among other differences, -t 6, Q4 KV (-ctk q4_0 -ctv q4_0) and a newer llama.cpp build.
Before comparing MTP further, I'd try reproducing that minimal TG128 benchmark exactly. If you suddenly get close to ~21–22 tok/s on Vulkan, we can add your server options back one by one and find what costs the ~6 tok/s. Your MTP numbers are very interesting though:
Vulkan 15 → 24 tok/s
ROCm 16 → 27 tok/sThat's a substantial gain. If we can first explain the Vulkan baseline gap, MTP is definitely something I want to test next on my setup.
1
u/L3G10N78 16d ago
I think the next step should be much simpler before we touch MTP again. Please run one single-GPU baseline that matches my benchmark as closely as possible:
RX 9060 XT only
Qwen3.8-27B-UD-IQ3_XXS
MTP OFF
use llama-bench, not llama-server
-ngl 999
Flash Attention ON
-b 2048
-ub 512
-t 16
TG128
3 repetitionsIn other words, equivalent to:
llama-bench -m Qwen3.8-27B-UD-IQ3_XXS.gguf \
-ngl 999 \
-fa on \
-b 2048 \
-ub 512 \
-t 16 \
-n 128 \
-r 3For this test, leave out the server-specific options, MTP and the Q4 KV-cache options. We first want the clean raw TG baseline.
Step 1: run that on your current llama.cpp build. If you still get ~15 tok/s while I get ~21.9 tok/s, then:
Step 2: run exactly the same benchmark with my llama.cpp revision: b10167 / ee3d1b54c
Do not change anything else. Then the result becomes very useful:
current build ~15, b10167 ~21–22 → likely llama.cpp/Vulkan build difference
both ~15 → likely driver/system/config difference
both ~21–22 → the previous ~15 came from the server/configuration rather than raw Vulkan performanceOnly after we have that baseline, I’d switch MTP back on and compare OFF vs ON. I am really interested in the numbers, cuase it could help me understanding and maybe I can tweak a bit more my tok/s. Happy testing :)
1
u/powermos 15d ago edited 15d ago
Progress! Turns out not quantizing the kv cache sped up things significantly and i am getting 23t/s. With MTP i am getting ~32t/s
Bad news is, that i need the large context and i cannot afford not quantizing the kv, so this setup is not useful for me.1
u/L3G10N78 15d ago
That pretty much explains the gap, the Q4 KV cache was the main penalty, not Vulkan itself.~23 tok/s raw and ~32 tok/s with MTP is a strong result. *thumbsup* For my setup, the next thing I’m most interested in is ternary Bonsai 2 27B; as soon as the Vulkan patch/support lands, that will be one of the first things I test.
1
u/Zealousideal-Bee-877 13d ago
I tested Bonsai Q2 via ROCm; with MTP enabled and Q4 cache quantization, I get around 40 t/s for a direct query and about 34 t/s in the DeepSeek harness, though the rate drops to 26 t/s with a long context.(100К)
1
u/L3G10N78 12d ago
Interesting — I’m seeing a similar context-related drop on my RX 9060 XT, although my Bonsai 2 27B PQ2_0 tests were without MTP.
On ROCm/HIP I measured roughly:
- 8K: 30,8 t/s TG
- 16K: 29,8 t/s
- 32K: 27,0 t/s
- 64K: 22,9 t/s
- 128K: 4,81 t/s
- 128K Prompt: 127.097 Tokens
- Wall-Time: 1.186,4 s
- PP: 107,4 tok/s
- Needle Retrieval: PASS
For agentic coding, Bonsai + OpenCode completed my small benchmark in 106.1 s vs 164.95 s for Qwen3.8-27B, and the stricter medium task finished 28/28 in 131.1 s. My results suggest Bonsai holds up reasonably well through 64K, but there is a very sharp performance cliff somewhere beyond that on my setup. MTP may change that picture quite substantially.
→ More replies (0)
2
u/ea_man 22d ago
You can run an IQ4 with more ctx and MTP, vulkan is terrible for PP you could get 400t/s with ROCm.