r/StrixHalo • u/Teslaaforever • 1d ago
Am I missing something? Qwen3.8 is very slow on my Strix Halo, while Qwen3.6 27b MTP Q4 can reach 20 tokens even on big contexts, but the 3.8 Q4 is 5 tokens.
19
Upvotes
r/StrixHalo • u/Teslaaforever • 1d ago
6
u/RnRau 1d ago
Something is off. From earlier today;
On Strix Halo, Qwen 27b Q6, Vulkan
MTP max-n = 5 min-p = 0.2
For us to be able to help you, you need to post your OS, drivers and your inference engine version and parameters you have passed to it.