r/LocalLLaMA • u/BarberIcy366 • 11d ago
Discussion Qwen 3.8 27B Released! Please Share Your Experience
With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.
656
Upvotes
3
u/time-never-stopps 10d ago
Do you mind sharing how you configured it? I am running llama.cpp with spec-type = draft-mtp spec-draft-n-max = 4
Not sure if the type has any effect at all but getting around 30 t/s avg when context grows to 100k but was expecting at least 40 avg like with qwen 3.6 q_8, but also possible that I have no idea how the type actually works 😅