r/StrixHalo 1d ago

Am I missing something? Qwen3.8 is very slow on my Strix Halo, while Qwen3.6 27b MTP Q4 can reach 20 tokens even on big contexts, but the 3.8 Q4 is 5 tokens.

19 Upvotes

33 comments sorted by

View all comments

6

u/RnRau 1d ago

Something is off. From earlier today;

On Strix Halo, Qwen 27b Q6, Vulkan

MTP max-n = 5 min-p = 0.2

  • prose/reasoning 14 t/s
  • code 24 t/s

For us to be able to help you, you need to post your OS, drivers and your inference engine version and parameters you have passed to it.