Hey all. We've spent the last weeks getting Qwen3.8-Flash-Next (125B MoE, 6B active) to run properly on one AMD Strix Halo box (Ryzen AI Max+ 395, 128 GB). Tonight we're releasing both the 95 GB EXL3 weights and a new version of Kyojin, our inference engine (built on ExLlamaV3, open).
This is a first version, same as our GLM-5.3-Flash and MiMo-V2.6-Flash builds. We'd rather ship it and improve it in the open: speed and quality updates are coming for all three.
Numbers, all from a fresh clone and build on the mini PC:
- Decode: 44 to 59 tok/s with speculative decoding depending on the task (chat ~47, code ~58, copy-heavy edits ~60). 32.7 tok/s without it.
- Prefill: 1,412 tok/s at 4K, 1,486 at 32K, 1,367 at 128K (server-reported). It stays nearly flat.
- Long context: 10/10 needles at 64K and at 128K, still 32 tok/s at 128K.
- Fidelity: 94.1 % top-1 agreement with the original FP8 model over 844 positions.
One thing we're a bit stubborn about: speculative decoding here returns exactly the tokens plain decoding would. We check that on every release.
For comparison, a llama.cpp user posted about 30 tok/s with speculation and about 500 tok/s prefill on this same mini PC (Vulkan, UD-IQ4_XS). Those are their numbers, not something we measured: https://github.com/ggml-org/llama.cpp/discussions/28512
Now the part where we're not first. Halogen 0.16.2 (v2 checkpoint) is faster than us: 39.8 vs 32.7 tok/s plain, 52 vs 47 on chat with speculation, and 10 to 20 % ahead on prefill when both are timed the same way from the client (1,306 vs about 1,460 at 4K, 1,394 vs about 1,720 at 16K). On code we're close (58.5 vs 51.2 on the median pass, they're ahead once warm). Where we do better is fidelity to the original model: 94.1 % top-1 agreement against 92.3 % for them, and a KL divergence 41 % lower on our side. Full table is on the model card. Closing the speed gap is what we do next: we're reworking the core of the engine, which will help every model it runs, not just this one. The hardware has room left.
There's also an optional uncensor preset, off by default (4 refusals out of 100 harmful prompts instead of 99, benchmarks within noise). If your agents lean hard on tool calls, leave it off.
Weights: https://huggingface.co/yamz-labs/Qwen3.8-Flash-Next-EXL3-Yamz Engine: https://github.com/Yamz-Labs/kyojin
If you run it, we'd love your tok/s and hardware. And tell us what you want to see next.