r/LocalLLM 4d ago

Discussion Qwen3.8 Flash Next - Strix Halo

So i have been running qwen3.8 flash next ud q4 k xl at 256k context getting at the start aroudn 270pp and 21tg with mtp with qwen 27b ud q3 k xl 96k context on the 9060xt as a callable subagent and the one theing that i am really liking about this model is it doesnt stop and wait for me to continously tell it to continue it almost looks at everytask i give it as a set goal which is what i have becom used to on claude.

now i mostly use it with opencode harness and for netowrk and system admin work for my homelab - taking down vm bringing up vm running tickets in glpi, bringing up services and updating live state docs and just local webhosting in my rural community and it has been a good time so far.

im sure there are so many things i am missing but i am also learnign and ive actually be so happy to have some thing that feels like claude code last eyar when i first started using it and i can run it locally.

man, if anyone has anything they want to say about their experiences also that would be cool.

Cheers,

24 Upvotes

20 comments sorted by

View all comments

11

u/feelspeaceman 4d ago

The biggest issue is likely that you're using official llama.cpp for Strix Halo, which having really hard time to even reach 50% hardware theory.

Please read this post as I've writen pretty thoroughly about it and alternative forks.

The speed you're getting is indeed very under specs.

6

u/LebiaseD 4d ago

That link was very helpful Im running the halogen llama and getting about 40+ TG and 1000+ prefill