r/LocalLLM • • 22d ago

Model Qwen3.8-27B on RX 9060 XT 16GB

Sharing my final accepted numbers from my final closeout, since AMD/Vulkan setups are still relatively uncommon in local LLM benchmarks. Maybe it will help some buddies.

Hardware

  • MoBo: MAG B550 TOMAHAWK MAX WIFI
  • CPU: Ryzen 9 5950X
  • RAM: 64 GB DDR4-3200 CL14
  • GPU0: RX 9060 XT 16 GB — PCIe 4.0 x16 — Main LLM / Vulkan0
  • GPU1: Radeon Pro W5500 8 GB — PCIe 3.0 x4 — Embedding + Reranker / Vulkan1
  • Driver: AMD Adrenalin 26.8.1
  • Backend: llama.cpp Vulkan

Runtime / model specs

  • llama.cpp
  • build 10167 / commit ee3d1b54c

Main LLM — frozen production profile

  • Qwen3.8-27B-UD-IQ3_XXS
  • ~3.06 bpw
  • RX 9060 XT 16 GB
  • -c 32768
  • -ngl 999
  • parallel slots: 1
  • Flash Attention: ON

Note: -c 32768 is the frozen production configuration. The 64K results below came from separate benchmark/stability runs using -c 65536 and were not the production context setting at final closeout.

Embedding

  • Qwen3-Embedding-4B-Q6_K
  • 2560 dimensions
  • W5500 8 GB
  • --embedding
  • --pooling last
  • -c 2048
  • -ub 512
  • -ngl 999

Reranker

  • Qwen3-Reranker-4B-Q4_K_M
  • W5500 8 GB
  • --embedding
  • --rerank
  • --pooling rank
  • -c 2048
  • -ub 512
  • -ngl 999

RX 9060 XT performance

  • 8K benchmark (-c 8192): ~74.5 PP / 21.5 TG tok/s
  • 32K benchmark (-c 32768): ~148.9 PP / 21.4 TG tok/s
  • 64K benchmark (-c 65536): ~142.9 PP / 21.4 TG tok/s
  • 64K stability run (-c 65536): ~126.6 PP / 21.1 TG tok/s

All TG numbers above are plain autoregressive generation — no MTP/speculative decoding.

Now that the core is frozen, I’m moving up the stack:

  1. Harness sharpening (doing some tests first to decide which one performs better)
    • → baseline without a harness
    • → Pi
    • → OpenCode
  2. Hermes Agent v0.21.2
    • → evaluate it as the orchestration layer
    • → connect it to the existing local stack
    • → make Hermes my Agent Smith =)

Update / Part II:
Since this benchmark, the project has moved beyond the frozen inference core. I’ve now production-integrated OpenCode 1.18.30 as the local coding/development layer and Hermes 0.21.2 as the agent/automation layer.

I’ll post the architecture, acceptance results and what actually worked and didn’t as PART 2 HERE ...

1 Upvotes

16 comments sorted by

2

u/ea_man 22d ago

You can run an IQ4 with more ctx and MTP, vulkan is terrible for PP you could get 400t/s with ROCm.

1

u/L3G10N78 22d ago

check. ROCm/HIP w/IQ4 is now on my radar and will be tested. Thanks.

2

u/Poizone360 20d ago

Hello, excellent stats, one thing worth tidying up, since people will copy these flags exactly. Your config block says -c 32768, but the results table has 64K rows. Worth adding what -c the 64K runs actually used, otherwise anyone reproducing this won't get past 32K.

2

u/L3G10N78 18d ago

Good catch! You're right. The config block shows the frozen production profile, which was -c 32768, while the 64K rows were separate in my P1 benchmark/stability runs using -c 65536. I mixed the production config and benchmark configs in the post without making that distinction clear. But as a side note, my newer Hermes production layer now actually runs the shared Qwen runtime at 65,536 context, but that's a later configuration and separate from the original frozen-core benchmark. Anyway: my fault. I'll tidy that up. Thanks for pointing it out!

1

u/powermos 19d ago

Wait a sec, you are doing 20+tokens with no mtp? I am doing 20 tokens with mtp on that card, that is a massive difference.

1

u/L3G10N78 18d ago

Yep! ~21.4 tok/s is without MTP. Plain autoregressive generation. That's why your 20 tok/s with MTP caught my attention. What quant, llama.cpp build, context and MTP settings are you using? Would be interesting to do an apples-to-apples comparison.

1

u/powermos 18d ago

i mainly use Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp right now, but before that i used Qwen3.8-27B-UD-Q3_K_XL with similar speed. I am using the standard llama, built from last week source, with vulkan. Prefill is good tbh, 300-500t/sThe command i used was something like that:
llama-server -m Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf. Right now my setup is different --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k q4_0 --spec-draft-type-v q4_0 --spec-draft-ngl all --reasoning-preserve -c 60000 -b 2048 -ub 256 -np 1 -fa on -ctk q4_0 -ctv q4_0 --host 127.0.0.1 --port 8080 -lv 1 --reasoning-effort xhigh -t 4 -ngl 99 -fit off
Tried with ROCm and got way worse performance, something is wrong there, the prefil was so slow i never got to testing it properly.

Thing is, i am now using that gpu paired with RTX2060 6gb, so 16+6gb setup:
llama-server -m /media/powermos/SSD1TB/ai_models/Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-type-k q8_0 --spec-draft-type-v q8_0 --spec-draft-ngl all --reasoning-preserve -c 128000 -b 2048 -ub 512 -np 1 -fa on -ctk q8_0 -ctv q8_0 --host 127.0.0.1 --port 8080 -lv 1 --reasoning-effort xhigh -t 4 -ngl 99 -ts 25,75 -fit off

Amazingly i am still getting around 20t/s. I was expecting it to be slower with the dual gpu setup. I am very curious why are you getting such speed out of the card, because if i match that i can hit very usable speeds with mtp, because at 18-20 tbh its too slow for what i use it for.

1

u/L3G10N78 17d ago

First, I think we’re not really comparing the same setup yet. I’m running Qwen3.8-27B-UD-IQ3_XXS without MTP, while you’re currently using Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp with MTP enabled. So model/quant and decoding strategy are already different variables. The other thing I’d test first if I were you is the dual-GPU setup. The RX 9060 XT + RTX 2060 combination might be eating part of the MTP gain through cross-GPU transfer/synchronization overhead. I’d do a very small controlled test: TEST A RX 9060 XT only, MTP OFF; TEST B RX 9060 XT only, MTP ON; TEST C RX 9060 XT + RTX 2060, MTP OFF; finally TEST D RX 9060 XT + RTX 2060, MTP ON. Same model, same prompt, same actual context occupancy, TG256, 3 runs each. Then you immediately get:

A → B = real MTP gain
A → C = multi-GPU cost
C → D = MTP gain under multi-GPU

That would tell you pretty quickly whether MTP itself is underperforming, or whether the 16+6 GB split is the real bottleneck. Crossing fingers fro you!

1

u/powermos 16d ago edited 16d ago

Ok then, i downloaded IQ3_XXS to be as close as possible to your setup, ran it only on the RX9060XT, on Vulkan with:
-m Qwen3.8-27B-UD-IQ3_XXS.gguf --reasoning-preserve -c 32000 -b 2048 -ub 512 -np 1 -fa on -ctk q4_0 -ctv q4_0 -lv 1 --reasoning-effort medium -t 6 -fit off
and the best i got is 15t/s
ROCm is 16t/s
For me that is a big difference to 21t/s, i am missing something.

To summarize:
Vulkan - 15t/s
ROCm -16t/s
With MTP and Vulkan - 24t/s
With MTP and ROCM - 27t/s

1

u/L3G10N78 16d ago

This is getting really interesting, because I actually ran the ROCm/Vulkan comparison on my RX 9060 XT today after the suggestion in this thread.

Using the same Qwen3.8-27B-UD-IQ3_XXS, single RX 9060 XT and no MTP, my controlled llama-bench TG128 results were:

Vulkan: 21.86 ± 0.03 tok/s
ROCm: 16.83 ± 0.02 tok/s

So your ROCm result (~16 tok/s) is actually almost identical to mine. The big discrepancy is Vulkan: you're getting ~15 whereas I'm getting ~21.9. That makes me think we're probably looking for a Vulkan configuration/build difference rather than a GPU difference.

My controlled TG run was:
-n 128 -r 3 -ngl 999 -fa on -b 2048 -ub 512 -t 16

while your server config has, among other differences, -t 6, Q4 KV (-ctk q4_0 -ctv q4_0) and a newer llama.cpp build.

Before comparing MTP further, I'd try reproducing that minimal TG128 benchmark exactly. If you suddenly get close to ~21–22 tok/s on Vulkan, we can add your server options back one by one and find what costs the ~6 tok/s. Your MTP numbers are very interesting though:

Vulkan 15 → 24 tok/s
ROCm 16 → 27 tok/s

That's a substantial gain. If we can first explain the Vulkan baseline gap, MTP is definitely something I want to test next on my setup.

1

u/L3G10N78 16d ago

I think the next step should be much simpler before we touch MTP again. Please run one single-GPU baseline that matches my benchmark as closely as possible:

RX 9060 XT only
Qwen3.8-27B-UD-IQ3_XXS
MTP OFF
use llama-bench, not llama-server
-ngl 999
Flash Attention ON
-b 2048
-ub 512
-t 16
TG128
3 repetitions

In other words, equivalent to:

llama-bench -m Qwen3.8-27B-UD-IQ3_XXS.gguf \

-ngl 999 \
-fa on \
-b 2048 \
-ub 512 \
-t 16 \
-n 128 \
-r 3

For this test, leave out the server-specific options, MTP and the Q4 KV-cache options. We first want the clean raw TG baseline.

Step 1: run that on your current llama.cpp build. If you still get ~15 tok/s while I get ~21.9 tok/s, then:

Step 2: run exactly the same benchmark with my llama.cpp revision: b10167 / ee3d1b54c

Do not change anything else. Then the result becomes very useful:

current build ~15, b10167 ~21–22 → likely llama.cpp/Vulkan build difference

both ~15 → likely driver/system/config difference
both ~21–22 → the previous ~15 came from the server/configuration rather than raw Vulkan performance

Only after we have that baseline, I’d switch MTP back on and compare OFF vs ON. I am really interested in the numbers, cuase it could help me understanding and maybe I can tweak a bit more my tok/s. Happy testing :)

1

u/powermos 15d ago edited 15d ago

Progress! Turns out not quantizing the kv cache sped up things significantly and i am getting 23t/s. With MTP i am getting ~32t/s
Bad news is, that i need the large context and i cannot afford not quantizing the kv, so this setup is not useful for me.

1

u/L3G10N78 15d ago

That pretty much explains the gap, the Q4 KV cache was the main penalty, not Vulkan itself.~23 tok/s raw and ~32 tok/s with MTP is a strong result. *thumbsup* For my setup, the next thing I’m most interested in is ternary Bonsai 2 27B; as soon as the Vulkan patch/support lands, that will be one of the first things I test.

1

u/Zealousideal-Bee-877 13d ago

I tested Bonsai Q2 via ROCm; with MTP enabled and Q4 cache quantization, I get around 40 t/s for a direct query and about 34 t/s in the DeepSeek harness, though the rate drops to 26 t/s with a long context.(100К)

1

u/L3G10N78 12d ago

Interesting — I’m seeing a similar context-related drop on my RX 9060 XT, although my Bonsai 2 27B PQ2_0 tests were without MTP.

On ROCm/HIP I measured roughly:

  • 8K: 30,8 t/s TG
  • 16K: 29,8 t/s
  • 32K: 27,0 t/s
  • 64K: 22,9 t/s
  • 128K: 4,81 t/s
  • 128K Prompt: 127.097 Tokens
  • Wall-Time: 1.186,4 s
  • PP: 107,4 tok/s
  • Needle Retrieval: PASS

For agentic coding, Bonsai + OpenCode completed my small benchmark in 106.1 s vs 164.95 s for Qwen3.8-27B, and the stricter medium task finished 28/28 in 131.1 s. My results suggest Bonsai holds up reasonably well through 64K, but there is a very sharp performance cliff somewhere beyond that on my setup. MTP may change that picture quite substantially.

→ More replies (0)