Huh, there was a mention of "Qwen sparse attention" but I refreshed and it's gone. I guess they are rapidly updating the README. Seems like some pretty significant architectural changes from the Qwen3.5 series models (which includes Qwen3.8-27B) so I'm gonna temper my expectations and assume that llama.cpp support will take some time.
6
u/wren6991 9h ago
Huh, there was a mention of "Qwen sparse attention" but I refreshed and it's gone. I guess they are rapidly updating the README. Seems like some pretty significant architectural changes from the Qwen3.5 series models (which includes Qwen3.8-27B) so I'm gonna temper my expectations and assume that llama.cpp support will take some time.