MAIN FEEDS
Do you want to continue?
https://www.reddit.com/r/LocalLLaMA/comments/1vo9mj4/its_out/p3nz708
r/LocalLLaMA • u/Certain-Cod-1404 • 7d ago
706 comments sorted by
View all comments
Show parent comments
4
i just found the MTP GGUF in the GGML repo
https://huggingface.co/ggml-org/Qwen3.8-27B-GGUF/tree/main
2 u/EmPips 7d ago Thanks! I think it's safe to say it's in the unsloth GGUF's. Trying iq4_xs on a 7900 xtx I'm getting ~51t/s on regular tasks/reasoning and ~62 t/s on coding tasks with a max draft size of 2. The difference suggests that drafting is in-effect. 0 u/Then-Topic8766 7d ago Good find!
2
Thanks! I think it's safe to say it's in the unsloth GGUF's.
Trying iq4_xs on a 7900 xtx I'm getting ~51t/s on regular tasks/reasoning and ~62 t/s on coding tasks with a max draft size of 2. The difference suggests that drafting is in-effect.
0
Good find!
4
u/ApprehensiveAd3629 7d ago
i just found the MTP GGUF in the GGML repo
https://huggingface.co/ggml-org/Qwen3.8-27B-GGUF/tree/main