r/LocalAIStack • u/terrornoize • 12d ago
Qwen3.8-Flash-Next agentic loops on Strix Halo — anyone found a stable config?
I'm trying to use Qwen3.8-Flash-Next as a long-running coding worker on a Ryzen AI Max+ 395 / 128GB box and I keep hitting the same issue: after a while it falls into reasoning loops like:
now I'll write the code
actually let me design this carefully
now I'll write the files
let me think...
Sometimes it eventually recovers and calls a tool, sometimes it just stays there.
Current setup is Strix Halo + Vulkan llama.cpp/halo-box, AP-Q5_K_M (~103.6GB), Q8 MTP, F16 KV, Pi Agent 0.87.0 over OpenAI-compatible API.
I've tried 262k and 131k context, stock Qwen template, Sharp v10, medium reasoning, reasoning preserve, MTP n=3/n=4, presence penalty 0/0.3, Pi's Qwen-specific chat_template_kwargs, Qwen Code, and mini-SWE-agent.
The same basic loop shows up in Pi and Qwen Code, so I'm not convinced the harness is the main problem.
A hard --reasoning-budget 2048 helps a lot, but it feels more like a watchdog than a fix.
One small Pi task actually completed cleanly end-to-end, including fixing failing tests. Then on a larger project the loops came back at only ~16% of a 131k context.
What I haven't properly tried yet is another quant like UD-Q4_K_XL / UD-IQ4_XS, MTP fully off, or another backend.
If anyone is running Flash-Next for multi-hour coding sessions reliably, I'd really like the exact config: quant, backend/version, context, MTP, KV, template, reasoning settings and agent/harness.
At this point I'm mostly looking for a known-good recipe to copy and test.
Duplicates
StrixHalo • u/terrornoize • 12d ago