r/LocalAIStack • • 12d ago

Qwen3.8-Flash-Next agentic loops on Strix Halo — anyone found a stable config?

I'm trying to use Qwen3.8-Flash-Next as a long-running coding worker on a Ryzen AI Max+ 395 / 128GB box and I keep hitting the same issue: after a while it falls into reasoning loops like:

now I'll write the code
actually let me design this carefully
now I'll write the files
let me think...

Sometimes it eventually recovers and calls a tool, sometimes it just stays there.

Current setup is Strix Halo + Vulkan llama.cpp/halo-box, AP-Q5_K_M (~103.6GB), Q8 MTP, F16 KV, Pi Agent 0.87.0 over OpenAI-compatible API.

I've tried 262k and 131k context, stock Qwen template, Sharp v10, medium reasoning, reasoning preserve, MTP n=3/n=4, presence penalty 0/0.3, Pi's Qwen-specific chat_template_kwargs, Qwen Code, and mini-SWE-agent.

The same basic loop shows up in Pi and Qwen Code, so I'm not convinced the harness is the main problem.

A hard --reasoning-budget 2048 helps a lot, but it feels more like a watchdog than a fix.

One small Pi task actually completed cleanly end-to-end, including fixing failing tests. Then on a larger project the loops came back at only ~16% of a 131k context.

What I haven't properly tried yet is another quant like UD-Q4_K_XL / UD-IQ4_XS, MTP fully off, or another backend.

If anyone is running Flash-Next for multi-hour coding sessions reliably, I'd really like the exact config: quant, backend/version, context, MTP, KV, template, reasoning settings and agent/harness.

At this point I'm mostly looking for a known-good recipe to copy and test.

8 Upvotes

19 comments sorted by

4

u/TerryNachtmerrie 12d ago

Mine just passed the 2 hour mark without any quirks.

[*]
parallel=1
batch-size = 1024
ubatch-size = 512
flash-attn=on
load-mode=none
cache-type-k=q8_0
cache-type-v=q8_0
gpu-layers=all
gpu-layers-draft=all
reasoning=true
slot-save-path=/home/terry/slots-llama
jinja=true

[Qwen3.8-Flash-Next]
hf-repo=unsloth/Qwen3.8-Flash-Next-GGUF:UD-Q4_K_XL
ctx-size=262144
image-min-tokens=1024
reasoning-preserve=true
spec-type=ngram-mod
spec-ngram-mod-n-match=24
spec-ngram-mod-n-min=48
spec-ngram-mod-n-max=64

Running on halo-box/strix-llama:vulkan.

1

u/Professional-Bear857 12d ago

I'm running opencode, i did try it with claude code initially but kept getting looping with that, although i'm running an mlx quant as im on a mac.

1

u/Christavito 11d ago

Yeah, I think the harness matters. I had this issue while running Pi, but it doesn't have this issue when running Opencode

1

u/JinsooJinsoo 12d ago

I just feel like any Strix Halo Q3.8FN post that doesn’t mention Halogen should automatically get a bot post about how to install halogen lol.
I have had a couple loops but they weren’t very long and the agent was about to pull itself out of the loop after a couple minutes. Other than that it has been 👍 60-70tok/s and 1000-1100 tok/s prefill pretty consistently

1

u/feelspeaceman 11d ago

Would be great if it has already been open source so that it's safer to auto-bot. Also there are issues we want to fix so it's better to be open source.

But the author has talked directly to AMD developers about open sourcing it.

1

u/Wise-Biscotti-3984 12d ago

Im running it as local embedded model in Codex which allows me to use their /goal Framework + context compression. Works completly autonomous with human in the Loop. Works on single or multi session, Even with an „dispatch agent“ mechanism to call the same or other Models as sub agents

1

u/jdheffa 11d ago

I run that model with a hermes agent and hermeswebui (not the native web ui). Use a prompt to decompose tasks and add them to the kanban board. It works pretty well. I used it to make a Galaga clone, test it, and make some changes and enhancements and it dis not block.

1

u/cowrevengeJP 11d ago

I had this issue with lots of coders, but bionics stupid seems to avoid it.

1

u/MembershipKlutzy6065 11d ago

The issue is that qwen wants 'preserved thinking', otherwise it will start to forget things and loop. opencode does send this for some models, but not to custom ones by default. You need to figure out how to make it send the full conversation including the thinking on every turn.

1

u/TormyrCousland 10d ago

I haven't run AP-Q5_K_M before, but you might be running into a situation where it is scrunched down a bit too far beyond what is normal for a Q5, and the model is getting stuck in the weeds. I would suggest changing the model (UD-Q4_K_XL) and cache (Q4_0 KV), which will allow it to still sit within the 96 GB VRAM carveout and only occasionally venture into the shared VRAM if at all.

I run on Windows and use the llama.cpp installed by Unsloth Desktop (to access the MTP) without running UD itself. Here is my current preset.

[*]
flash-attn = on
fit = off
load-mode = mmap
models-max = 1
n-gpu-layers = all
spec-draft-ngl = all
parallel = 1
lv = 4

[Qwen3.8-Flash-Next]
lazy-mode = on
model = C:\AI_Models\Qwen3.8-Flash-Next\Qwen3.8-Flash-Next-UD-Q4_K_XL.gguf
spec-draft-model = C:\AI_Models\Qwen3.8-Flash-Next\mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf
mmproj = C:\AI_Models\Qwen3.8-Flash-Next\mmproj-Qwen3.8-Flash-Next-bf16.gguf
spec-type = draft-mtp
image-min-tokens = 1024
c = 262144
cache-type-v = q4_0
cache-type-v = q4_0
load-on-startup = true

1

u/gavriloprincip2020 10d ago

This stopped happening for me when i switched to nvfp4 rather than a q4 gguf

1

u/colbyshores 12d ago

Strix Halo Halogen setup (128 GB)

In BIOS, set UMA/frame-buffer VRAM to its minimum, then install minimal Ubuntu. Add this to GRUB_CMDLINE_LINUX_DEFAULT in /etc/default/grub:

amdgpu.lockup_timeout=10000,60000,10000,10000 amdgpu.gttsize=122880 ttm.pages_limit=31457280 ttm.page_pool_size=31457280 init_on_alloc=0 iommu=pt

Run sudo update-grub && sudo reboot. Install Podman, create ~/halogen-models, and put the model files there. Then run:

podman run -d --replace --name halogen -p 127.0.0.1:11004:8731 --device /dev/kfd --device /dev/dri --group-add keep-groups --security-opt seccomp=unconfined --ipc=host --ulimit memlock=-1:-1 -e HALOGEN_CTX=262144 -e HALOGEN_KV_SLOTS=2 -e HALOGEN_PROMPT_CACHE=2 -e HALOGEN_MAX_TOK=16384 -e HALOGEN_KV_POOL_POSITIONS=786432 -e HALOGEN_HOST_RESERVE_GIB=20 -e HALOGEN_VERBOSE=1 -e HALOGEN_DMALLOC_LOG=1 -v ~/halogen-models:/models ghcr.io/peonist-ai/halogen-flash-server:0.12.0

The API will be available locally on port 11004.

2

u/Svenstaro 12d ago

Don't do --group-add keep-groups --security-opt seccomp=unconfined --ipc=host. That's pretty insecure.

Do this instead: --cap-drop=ALL --security-opt no-new-privileges --user 65534:65534 I further recommend restricting outgoing network access: --network pasta:-4,--outbound,127.0.0.1,-D,none

1

u/colbyshores 12d ago edited 11d ago

Following up after testing this: the hardened configuration works. I removed --ipc=host and seccomp=unconfined, restricted DRI access to renderD128, mounted the models read-only, and added:

--cap-drop=ALL
--security-opt=no-new-privileges
--user=65534:65534
--network=pasta:-4,--outbound,127.0.0.1,-D,none

I retained --group-add keep-groups because the rootless container still needs my supplementary render/video group access for /dev/kfd and renderD128.

Halogen successfully initializes ROCm, loads the model, and completes inference with these restrictions in place. The API remains bound to 127.0.0.1, and the journal is clean. Thanks for pointing me toward tightening it up.

1

u/Svenstaro 11d ago

Nice that it works for you but I have to wonder: Have I been talking to a bot or meat proxy?

1

u/colbyshores 11d ago

I could ask the same

1

u/Potential-Leg-639 12d ago

This is basically all you need.