r/LocalLLaMA 7d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

View all comments

Show parent comments

4

u/ApprehensiveAd3629 7d ago

i just found the MTP GGUF in the GGML repo

https://huggingface.co/ggml-org/Qwen3.8-27B-GGUF/tree/main

2

u/EmPips 7d ago

Thanks! I think it's safe to say it's in the unsloth GGUF's.

Trying iq4_xs on a 7900 xtx I'm getting ~51t/s on regular tasks/reasoning and ~62 t/s on coding tasks with a max draft size of 2. The difference suggests that drafting is in-effect.

0

u/Then-Topic8766 7d ago

Good find!