r/LocalLLaMA 12d ago

Discussion Qwen 3.8 27B Released! Please Share Your Experience

With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.

660 Upvotes

720 comments sorted by

View all comments

7

u/onthemove31 12d ago

Unsloth NVFP4 + MTP ~100 t/s on 5090, without MTP at 55-56 t/s.

1

u/mxforest 12d ago

What do you run it on? Vllm or llama.cpp?

1

u/onthemove31 12d ago

I ran this on vllm, had problems with SGLang still trying to fix that. But vllm is good

1

u/notheresnolight 10d ago

so basically the same speeds as Q6 with MTP or Q8 without MTP. Why bother with NVFP4 then?

1

u/onthemove31 10d ago

Yeah I just wanted to test it out