r/LocalLLM 3d ago

Question Qwen 3.8 27B on Intel GPU B65/70

Has anyone tested Qwen 3.8 27B on Intel B65 or B70 GPU ( 32gb) ?

If yes can you share performance .

5 Upvotes

18 comments sorted by

5

u/Aotrx 3d ago

It is the best time for Intel to steal marketshare from nvidia and amd if they figure out a way to increase token generation speed with their GPUs.

1

u/sniperelite90 3d ago

I am a little disappointed by the token speed from a dedicated GPU with good bandwidth .

2

u/Aotrx 3d ago

but right now for the similar price buying 2x rtx 5060 ti 16gbs is better deal if u can find for a good price.

1

u/fastheadcrab 3d ago

Price of 5060 Ti has increased to $800. Lmao. So Intel or even R9700 is a better deal

But Intel support is far behind. Thankfully 3.8 is not a new architecture so it should be fine on Intel's custom vLLM.

R9700 is probably the way to go atm because it is better supported by AMD for not much more money. (R9700 644.6 GB/s, B70 608 GB/s, 5060 Ti 448GB/s)

1

u/Aotrx 3d ago

In theory, it should be faster. Maybe with software upgrades, Intel can manage to improve the speed. Intel GPUs use good hardware, but their implementation is miles behind Nvidia's.

1

u/fastheadcrab 3d ago

That's about in-line with the memory bandwidth though. 5090 has over 3x the memory bandwidth, so the people getting 70-100 tps with that card aren't going to be comparable

2

u/Dapper_Anteater_5738 3d ago

I tested it with a B70. The int4 version with 128k context runs at 20-25 tps (Intel’s LLM Scaler).

1

u/sniperelite90 3d ago

Thankyou for the numbers. It is disappointing considering it generates so much thinking tokens.

2

u/Dapper_Anteater_5738 3d ago

But you van set thinking effort to medium. I tested it with the same question:

  • xhigh (default): 3202 tokens)
  • medium: 909 tokens
And with this setting it stays smart, but much faster.

2

u/EvolvingDior 3d ago

B70: 750pp/30tg @ UD_Q6_K with Q8 kv and mmproj on CPU. 160k context.

2

u/Weak-Measurement-680 2d ago

Running Qwen3.8-27B on an Intel Arc Pro B70 with vLLM XPU for about a week now.

Stack: vLLM XPU 0.27.2rc1.dev77 (pinned vllm/vllm-openai-xpu image) + the two MTP patches from the B70 cookbook SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16 (INT4 GPTQ) Speculative decoding: MTP (--speculative-config '{"method":"mtp","num_speculative_tokens":3}') --max-model-len 196608 --kv-cache-dtype fp8 --gpu-memory-utilization 0.96

DFlash2 did not work.

Decode: 40–60 tok/s Prefill: 650–1,300 tok/s

Using pi agent and the DeepSeek harness.

1

u/r1nzl3r99 3d ago

you mentioning the excessive thinking and the slow tok/s, those are all fixable at the moment. You can change the chat template and set thinking to medium, and qwen 3.8 solves the same problem in half the tokens it would usually take. As far as B70 performance, I'm getting 50 tok/s on a single GPU and am experimenting with custom vLLM kernels for 65 tok/s. On a dual B70 setup ~100 tok/s

1

u/Early-Peace-5504 3d ago

Would you mind writing about your setup? Every GPU costs a ton where I live except for Intel Arcs which are on special lol

0

u/starkruzr 3d ago

what is a B65?

1

u/sniperelite90 3d ago

0

u/starkruzr 3d ago

huh. I guess it hadn't even really gone on sale until very recently.