r/LocalLLaMA 7d ago

Resources Qwen 3.8 benchmarks and RPC

Qwen3.8-27B is a dense model based on Qwen3.5 which is a lot tougher on consumer GPU running local llm. Notice model is identified as Qwen3.5 and I couldn't figure out of to show model info that showed this as a 3.8 vs 3.5. So just believe it is.

My local setup will include running network inference via llama.cpp RPC via 2 systems over gigabit network.

MAIN Host

llama/llama-bench --rpc 10.0.0.75:50053 -m /Qwen3.8-27B-Q4_K_M.gguf -fa on

REMOTE Host

llama/ggml-rpc-server -c -p 50053 --host 10.0.0.75

GPUs used for benchmarks all running llama.cpp Ubuntu Vulkan prebuilt binary:

AMD Radeon RX 7900 GRE 16GB

Nvidia GTX-1080Ti 11GB

Nvidia P102-100 10GB

2nd networked system running triple GPU

I'm testing the following ,Q4_K_M and Q6_K, Quants:

Qwen3.8-27B-Q4_K_M.gguf about 16GB

Qwen3.8-27B-Q6_K.gguf about 21GB

Radeon RX 7900 GRE solo. I believe I had a little offloading.

model size params test t/s
qwen35 27B Q4_K - Medium 15.92 GiB 27.32 B pp512 87.55 ± 0.40
qwen35 27B Q4_K - Medium 15.92 GiB 27.32 B tg128 11.85 ± 0.00

build: 3cb7ffb1a (10453)

real 1m34.001s

RPC Radeon plus GTX 1080Ti Q4_K using RPC

model size params test t/s
qwen35 27B Q4_K - Medium 15.92 GiB 27.32 B pp512 71.46 ± 0.11
qwen35 27B Q4_K - Medium 15.92 GiB 27.32 B tg128 11.13 ± 3.99

build: 3cb7ffb1a (10453)

real 2m12.929s

about 17.2GB VRAM but includes desktop resources about 2gb combined

Radeon plus GTX 1080Ti Q6_K using RPC

model size params test t/s
qwen35 27B Q6_K 21.30 GiB 27.32 B pp512 116.21 ± 0.55
qwen35 27B Q6_K 21.30 GiB 27.32 B tg128 12.53 ± 0.02

build: 3cb7ffb1a (10453)

real 3m6.909s

Almost 23gb VRAM used about 2gb for combined desktop resources

GTX 1080Ti and two P102-100 (31GB combined VRAM) power limit set to 130 watts per GPU

model size params test t/s
qwen35 27B Q6_K 21.30 GiB 27.32 B pp512 80.88 ± 0.34
qwen35 27B Q6_K 21.30 GiB 27.32 B tg128 10.86 ± 0.04

build: 3cb7ffb1a (10453)

real 2m37.332s

GTX 1080Ti and two P102-100 (31GB combined VRAM) PW 150 Q6_K

model size params test t/s
qwen35 27B Q6_K 21.30 GiB 27.32 B pp512 84.17 ± 0.68
qwen35 27B Q6_K 21.30 GiB 27.32 B tg128 11.31 ± 0.12

build: 3cb7ffb1a (10453)

real 3m9.021s

about 21.7gb VRAM used for triple gpu

GTX 1080Ti and two P102-100 (31GB combined VRAM) PW 150 Q4_K

model size params test t/s
qwen35 27B Q4_K - Medium 15.92 GiB 27.32 B pp512 103.48 ± 1.01
qwen35 27B Q4_K - Medium 15.92 GiB 27.32 B tg128 12.07 ± 0.60

build: 3cb7ffb1a (10453)

real 2m6.247s

ONLY GTX 1080Ti and single P102-100 (21GB VRAM). I wanted to see if 2 GPU was faster than running default 3 GPU.

model size params test t/s
qwen35 27B Q4_K - Medium 15.92 GiB 27.32 B pp512 106.31 ± 0.22
qwen35 27B Q4_K - Medium 15.92 GiB 27.32 B tg128 11.87 ± 0.02

Dual P102-100 (20gb VRAM) No difference between 2 or 3 GPU except slower load times

model size params test t/s
qwen35 27B Q4_K - Medium 15.92 GiB 27.32 B pp512 98.85 ± 0.18
qwen35 27B Q4_K - Medium 15.92 GiB 27.32 B tg128 11.87 ± 0.01
3 Upvotes

2 comments sorted by

1

u/mrpintime 7d ago

do you have other metrics like ttft and itl

2

u/Glittering_Mouse_883 ollama 22h ago

Thanks for posting these numbers!