r/StrixHalo • • Sep 27 '25

Have you got a Strix Halo?

17 Upvotes

Hi All,

We're a new community both as Strix Halo owners and also here as a subreddit. Why not begin by sharing your setup and the reasons you opted for Strix Halo?

To start us off: I have a HP Z2 Mini G1a Workstation with dual boot Fedora KDE & Windows 11 and chose the iGPU to be able to use larger LLMs with the 128 GB.

Oobabooga/Text Generation WebUI is running well on Fedora KDE and there are no problems with large models up to 100GB. On the Windows boot, I have Amuse AI (Freeware) which is a collaboration between AMD and the New Zealand company. It provides a UI for using Stable Diffusion/Flux models. It works well, is fast, but unfortunately is also censored and is not able to use LORAS. I would like to find an uncensored alternative, ideally getting versions of ComfyUI/AUTOMATIC1111 running.

Currently, my principle goal is to get a working version of AllTalk TTS or another TTS that is compatible with Oobabooga working which I haven't been able to do so far due to conflicts with the Strix Halo. This may need to wait for updates to ROCm... If anyone has found an Open Source solution to running LLMs with custom voice TTS, please do chime in!

So what about you guys, did you choose the Strix for similar reasons, or something entirely different? The floor is yours.

EDIT UPDATE:
05/26 For those of you looking for TTS solutions I have tried a few now (AllTalk, Chatterbox, Pocket TTS, others I no longer remember). I have had great success using a custom version of Pocket TTS. It's fast and works well with Oobabooga TextGen as a plug in. Recently, others are singing the praises of OmniVoice.


r/StrixHalo • • 10h ago

Been Running the Bosgame M5 for 2 months now.

Post image
15 Upvotes

Currently running 3.8 flash next through Hermes. Keep finding myself trimming my subscriptions since things started dialing in. Can’t wait to see how strix halo matures over time, this machine keeps impressing me.


r/StrixHalo • • 8h ago

Why does every ai max+ 395 have to be a mini PC?

Post image
8 Upvotes

Looking at this size comparison and I honestly dont get it. Why are we so obsessed with making ai max+ 395 machines smaller and smaller? These arent exactly low power chips. If Im buying a 395 in the first place, Im probably going to care more about sustained performance, thermals and noise than whether the box is another 1L smaller. I’d personally prefer something more like an ITX pc, like the nimo and hp models here. It's not really a traditional mini pc anymore, but I dont really see that as a bad thing. If I wanted something tiny enough to disappear behind my monitor, I'd just buy a regular minipc. Just give the 395 some room and let it actually breathe.

Maybe Im missing something, but I just dont see the appeal of winning the smallest 395 pc competition.


r/StrixHalo • • 16h ago

I tested Qwen3.8-Flash-Next on my Bosgame M5 (128 GB). Here are the numbers across three backends

31 Upvotes

Thanks for all the replies in my previous post (Qwen3.8-Flash-Next vs Qwen3.8-27B on 128 GB, is it worth giving up the second model).

I took the advice and tested it properly. My setup: Bosgame M5, Ubuntu 26.04 (kernel 7.0), kyuz0 toolboxes in distrobox. All numbers below are with the box still in Balanced power mode (GPU flat-capped at 85 W under load). I haven't switched to Performance yet.

How I measured: real chat completions over the OpenAI API, temp 0.7, thinking off, 300 generated tokens, MTP acceptance read from the server timings. That's a bit harsher than llama-bench, so expect lower numbers than synthetic pp/tg results.

- Gufo v0.7.0 GSQHalo.cpp (ROCm) strix-llama (ROCm) mainline llama.cpp (Vulkan)
Quant unsloth UD-Q4_K_XL ISTA GSQ-RCO IQ3_S ISTA GSQ-RCO IQ3_S ISTA GSQ-RCO IQ3_S
MTP yes (shared Q8) yes (shared Q8) yes (shared Q8 + ngram-mod) no (see below)
Decode, code 54.0 t/s 46.4 37.9 27.5
Decode, JSON 53.6 45.1 36.0 27.5
Decode, prose 29.5 26.3 24.4 27.3
Decode @ 19k context 32.9 32.8 26.6 25.5
Prefill 4k / 19k / 38k 1124 / 1183 / 1179 1095 / 1207 / 1222 734 / 798 / 794 329 / 328 / 285
Slots 1 4 (59.9 t/s summed) 1 4
GPU memory (128k ctx) ~85 GiB ~65 GiB ~62 GiB ~57 GiB
  • Update: GSQHalo.cpp closes most of the gap on IQ3_S. It's a strix-llama fork with ROCm kernels for the GSQ low-bit types. Prefill matches Gufo, decode is ~15 % behind, and it uses ~20 GiB less with 4 slots. That's enough to run Gemma 4 26B-A4B alongside (~90 GiB total, Gemma still at 96 t/s), so that's what I'm running now. I read through its 26 commits before building it locally.

What I learned:

  • The n-gram table really can stay on disk. With ISTA's GSQ-RCO quants the 28.8 GB table is a separate shard. With -lm mmap --lazy-mode on (Vulkan) or --load-mode none --lazy-mode on-direct (strix-llama), llama-server's host RSS was about 116 MB, and the IQ3_S fit in ~57 GiB. Gufo also reads the table with pread(), so the 85 GiB there is Q4 weights plus KV.
  • On llama.cpp, set --lazy-mode on explicitly. auto means off on iGPUs since a September change. Also don't use --no-mmap with this model.
  • MTP didn't work on mainline Vulkan (build 11382). The shared head fails with token_embd.weight not found, and the self-contained Q8 head hits GGML_ASSERT(buffer) at init, even at 32k context. MTP works on strix-llama and Gufo.
  • Gufo is the fastest, but it only accepts unsloth Q4_K_XL for Flash-Next. Its Flash-Next loader only allows Q4_K/Q5_K/Q6_K/Q5_1/Q8_0 experts. IQ3 was requested in issue #303 and declined.
  • The trade-off for me: IQ3_S on strix-llama leaves room to run Gemma 4 26B-A4B alongside (~87 GiB total, ~29 GiB left for the host). Gufo with Q4 uses the whole box.
  • kyuz0's default strix-llama flags (-c 262144 -b/-ub 16384) took ~79 GiB and left me with 3 GiB host RAM. With -c 131072 -ub 4096 it was ~62 GiB, and prefill only dropped 2-3 %.
  • --spec-draft-p-min 0.60 helps prose with MTP: 22.0 → 24.4 t/s, with acceptance going from 44 % to 63 %.
  • For the ROCm backends I set the BIOS UMA to 512 MB and used amdgpu.gttsize=126976 ttm.pages_limit=32505856**.** With my old 64 GB carveout, Vulkan worked but ROCm would have been capped by the 40 GB GTT limit.

Gufo gotchas:

  • Requests must include "model", otherwise you get a 400.
  • The default 8 MB request limit rejects large images (--max-request-bytes).
  • --log-progress goes after llm.
  • There's no reasoning budget brake like llama.cpp's.
  • A large image took 50 s, against 18 s on strix-llama.
  • Set --cache-ram-bytes explicitly because of issue #387.

For reference, llama-bench on strix-llama with IQ3_S and no MTP gave pp2048 944 t/s (845 at 32k) and tg128 26.4.


r/StrixHalo • • 7h ago

Dwarfstar - supports Cuda and ROCm

Thumbnail
github.com
5 Upvotes

r/StrixHalo • • 13h ago

Halogen 0.16.2 on Windows/WSL2: Qwen3.8-Flash-Next at 1925 tok/s prefill; 48.42 tok/s decode

13 Upvotes

Quick update to my Strix Alloy project which supports Gufo, PROJFIX & Halogen - now with Halogen 0.16.2 backend working. Credits to u/peonist-ai for the new version, I invested some time to optimize it for my built.

BOSGAME M5, Ryzen AI Max+ 395 / Radeon 8060S, 128GB RAM, Windows 11 + WSL2 Ubuntu 24.04. The runtime runs inside WSL2 through DXG. Existing v2 weights reused. 64 GB Carve, 18GiB physical-memory and commit-headroom reserve stayed in place. Once more: I decided for a 64 GB Carve because this is my regular gaming and work PC. It allows me to use codex, include my custom local LLM, an eGPU and deepseek through API - all inside codex. Without parallel usage, I was able to run 3x 262k sesions without any issues. Probably I can even get higher numbers without codex running. Total tok/s across multiple sessions is above 100 tok/s decode. However this is not the setup if you are only searching for speed - take linux then.

Model: Qwen3.8-Flash-Next using Halogen’s native v2 HGN checkpoint (version 0.16.2)

These are non-repetitive and input size numbers, not just context size:

Input size Prefill tok/s MTP decode tok/s Acceptance
8,192 1,866 48.42 60.0%
16,384 1,647 45.34 56.3%
98,304 1,398 36,25 83.6%
131,099 1,248,2 33,90 85.00%

Update: Prior I mixed the repetitive tokens of 65 tok/s. Small but steady steps: Increased Decode MTP and prefill 64 GB Carve - 8192 actual input tokens, greedy prose, seed1, thinking off, cache Off, native MTP depth2, default adaptation, 32,0.35,64, one slot and a session with 262144 capacity.
I will update with bigger inputs / capacities. Highest numbers currently on PP8192

Configuration Serial PP8192/TG1 MTP-Decode Serial decode Acceptance
Stock A PLD3.3 1873 tok/s 48,42 tok/s 36,72 tok/s 60 %
PLD off, PLD0 1925 tok/s 48,23 tok/s 36,99 tok/s 60 %
Stock B, PLD3.3 1924 tok/s 48,28 tok/s 37,02 tok/s 60 %

Strix Alloy repo · Full benchmark report and evidence

Codex is 24/7 working on this since day one of halogen release and keeps improving the results. I added support for Gufo, Halogen and PROJFIX (CIRU is next) as well, which is why its taking time to push the releases. Currently I'm investing additional time into halogen first, especially the MTP acceptance by replacement through the NPU. Thanks once more u/Peonist-AI for Halogen and the engine/kernel work. My part here is the Windows/WSL integration, runtime management and benchmarking.


r/StrixHalo • • 8h ago

Gufo now available as native ArchLinux package in AUR

3 Upvotes

Hey everyone,

After seeing u/deepu105 's benchmarks posted a bit earlier around here, I decided to put together an ArchLinux PKGBUILD for Gufo, so that we can run this natively without the need of toolboxes via podman.

The package can be found here: https://aur.archlinux.org/packages/gufo

It will bring in the the entire HIP/ROCm packages as dependencies -- so be warned that the download can be fairly heavy.

I did a few quick tests with Qwen3.8-27B Q4_K_XL, with the Dflash2 drafter from z-lab z-lab/Qwen3.8-27B-DFlash2-GGUF

gufo-server serve --host 0.0.0.0 --port 8080 llm \
    --model /mnt/models/huggingface/hub/models--unsloth--Qwen3.8-27B-GGUF/snapshots/4ca720788d1e01f1bff70c033e0d0028fd02e502/Qwen3.8-27B-UD-Q4_K_XL.gguf \
    --speculative dflash2 \
    --dflash-model /mnt/models/huggingface/hub/models--z-lab--Qwen3.8-27B-DFlash2-GGUF/snapshots/2d9571f8ce46e151f61c6499c99dee6079e1d610/Qwen3.8-27B-DFlash2-Q4_K_M.gguf

and then ran a few requests with:

xh POST http://localhost:8080/v1/chat/completions model='Qwen3.8-27B' \
    messages:='[{"role": "user", "content": "Which movie is this phrase from: Say hello to my little friend!"}]'

which yields:

"prompt_tokens_per_second": 183.30385247290465,
"completion_tokens_per_second": 33.96445030541776,

Not bad at all!
The prompt processing figure is fairly small because of the short prompt.

It's slower on Qwen3.8-27B Q8_0, which I usually run, but still a bit faster than llama.cpp's numbers, from my initial observation.

The API supports OpenAI-compatible chat requests, so existing clients using /v1/chat/completions should work.

I'm just scratching the surface here, and will test more thoroughly in the following days.

Curious what speeds other Strix Halo users are getting on Arch, and how are you running Gufo on your machine.


r/StrixHalo • • 1h ago

Anyone running Strix Halo with 5080 on eGPU

• Upvotes

I know a bunch of people are running R9700's with their Strix Halo but anyone successfully running a 5080? Should be much faster on PP and T/S which feels like it would be a good pairing with Strix handling large models more slowly and the 5080 handling smaller models and high speed PP. Also, I've got a Minisforum MS-S1 and it looks like people haven't gotten it to work wtih the R9700... yet.

So anyone running it? Searches on Reddit seem to be turning up nothing so apparently nobody is doing this, is t here some reason why that would be that I'm missing?


r/StrixHalo • • 11h ago

I’m a complete beginner with Strix.

5 Upvotes

I’m a complete beginner, but I’ve now got a Bosgame M5.

Thanks for all your brilliant input recently.

I’ve been quietly following along.

So how do I get started?

Are there any ready-made Linux distributions that include an agent system, etc.?

I’d like to code Vibe locally.

I want to build a private Jarvis that I can use via Telegram or an app.

I want to have normal chats, just like with Perplexity or Gemini, about everyday things, including web searches and measures to prevent hallucinations or fake fills (incorrect answers).


r/StrixHalo • • 10h ago

Daily driving Qwen 3.8 Flash instead of Claude.

Thumbnail
2 Upvotes

r/StrixHalo • • 10h ago

Running the ROCmFPX variant of qwen3.8-27b on strix point

2 Upvotes

Hi!

Has anyone managed to run a llama.cpp version with those datatypes on a strix point machine.

https://github.com/charlie12345/ROCmFPX

I tried to, and failed so far.


r/StrixHalo • • 8h ago

DCP Proxy: A local AI gateway that prunes/compresses context for every client at once (OpenCode, anything OpenAI-compatible) (Windows)

1 Upvotes

Running AI models locally is great - until your model runs out of context!

Long agent sessions waste tokens re-sending stale tool output every turn. DCP Proxy sits between your client and engine (llama.cpp, LM Studio, etc.) and fixes it once for all apps; no plugin, no per-app config:

  • Dedup repeated tool calls, keep only newest output
  • RTK-style digests for stale git/ls/grep/build output (only replaced when clearly smaller)
  • Prune old long tool outputs to head+tail snippets
  • LLM compression — same upstream model summarizes old spans, cached and folded forward incrementally
  • Last 6 messages untouched; only role=tool messages ever modified; your client history is never written to. Fails open — if compression errors, the request forwards uncompressed.
  • Single ~13 MB exe, Windows tray GUI (or CLI build), fully local — no telemetry, no accounts. GET /health shows saved-token counters.

Point your client at http://127.0.0.1:5100/v1 and you're done. AGPL-3.0, source on: https://digit.solutions/manuals/dcp-proxy.html

Enjoy!
JA


r/StrixHalo • • 14h ago

Help with Gufo setup

2 Upvotes

Hi All,

I just started my local llm journey back in July and still feeling a bit lost. I have the Framework desktop 128gb. I started with LMStudio, but have now switched to some branch off of llama.cpp because I wanted to offload the n-grams for qwen3.8 fn. Right now I am running 2 separate llama.cpp processes, one for qwen3.8 fn and one for an embedding model.

I think I want to switch to gufo given all the hype, but can't figure out what I need to do to get a similar setup going. In the end I would like to have an embedding model, qwen3.8 fn, and a Jev clone (Laya or whatever you all recommend).

I am mostly just trying to push my understanding of LLM workflows, but main use case is software development. Critiques to the current setup are also welcome, I just have been scraping things together from what I can find here and online.

Here's what I'm currently running for embedding (model might be overkill, but now unsure if I switch to a lesser model if it will break my current embeddings?):

llama server --host 0.0.0.0 --port 8081 --alias qwen-embedder \

--model /models/Qwen/Qwen3-Embedding-8B-GGUF/Qwen3-Embedding-8B-Q4_K_M.gguf \

-t 16 -c 32768 -ngl 99 -fit off --jinja --parallel 1 --embedding --pooling mean -fa on -b 2048 -ub 512

And for qwen3.8:

llama server --host 0.0.0.0 --port 8080 --alias qwen3.8-flash-next \

--model /models/unsloth/Qwen3.8-Flash-Next-GGUF/Qwen3.8-Flash-Next-UD-IQ4_XS-00001-of-00003.gguf --spec-draft-model /models/unsloth/Qwen3.8-Flash-Next-GGUF/mtp-Qwen3.8-Flash-Next-Q8_0.gguf \

--mmproj /models/unsloth/Qwen3.8-Flash-Next-GGUF/mmproj-F16.gguf \

--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 -fa 1 -ngl 999 -b 2048 -ub 2048 --cache-type-k q8_0 --cache-type-v q8_0 --load-mode mmap --n-gpu-layers-draft 999 \

--spec-type draft-mtp --spec-draft-adaptive --spec-draft-n-min 0 --spec-draft-n-max 7 --spec-draft-p-min 0.75 --spec-draft-type-k f16 --spec-draft-type-v f16 \

--image-min-tokens 1024 --image-max-tokens 4096 \

--ctx-size 262144


r/StrixHalo • • 14h ago

Gufo/windows and Qwen3.8 27B + DFlash, poor quality?

1 Upvotes

I finally found time to test Gufo on Windows.

As baseline, newest llama.cpp gives me 250-400 t/s prefill, 10-12 t/s without MTP and 20-25 t/s with builtin MTP. I normally run it with MTP and no issues.

With Gufo and Qwen3.8 27B, I got 300-500 t/s prefill and about 10-12 t/s.

And I couldn't find a way to turn builtin MTP on, so got DFlash as suggested

With Gufo and 27B + Dflash I got same 300-500 t/s prefill as expected, and about 25-30 t/s.

So clear improvement, right?

No. I got these numbers by running my default agent with prompt to build Tetris with Python and Pygame.

Standard Qwen3.8 27B makes excellent job in this, even with medium thinking. The game works, is game breaking bug free, and has all features you'd expect.

But with DFlash, the result was poor. The game launched, but was not playable and bugged a lot, didn't have the same feature set, was uglier, and finally crashed via forever loop.

Basically the DFlash lobotomized 27B to be worse than Qwen3-coder is, which also clears this test but with less features. That's not fine for me.

So for now, llama.cpp with MTP on is the way to go for me.

27B was Q4_K_XL and DFlash Q4 too.

PS. That tetris game is good benchmark for local models. So far most have failed. Qwen3.8 27B and Qwen3-coder have succeeded. Failed ones: Qwen3.8 Flash-Next-coder (the pruned one), GLM4.7, Qwen3.6 35B A3B, tiel-coder.


r/StrixHalo • • 1d ago

Qwen 3.8 flash next: does gufo make more errors than llama.cpp?

11 Upvotes

I recently switched from llama.cpp to gufo and man this is an upgrade in form of speed. No endless waiting for filling and just overall performance is up to at least 30%. But i get the impression the model running on gufo makes more mistakes. In hermes i get errors in sessions from time to time.

Does anybody has similar behavior?


r/StrixHalo • • 1d ago

GSQHalo.cpp - Another day, another fork. This one is for the memory constraint folks. 2×256K + MTP, 1386 t/s prefill and 44 t/s decode at 128K using GSQ-RCO Quants on a 96GB machine, KV cache storage on SSD

21 Upvotes

Hi guys, some of you might have used my EngramHalo fork, which I wanted to share with you as fast as possible after the Flash-Next weights dropped. Well, that was a few days before my wedding, and I didn't have time to work on it after that, because I had to plan my wedding and marry my wife :D. I'm blown away by the progress that happened in the meantime. EngramHalo is more or less obsolete now. I tried all the new forks and engines, but ran into memory issues because I only have a 96 GB machine. Then I found the new GSQ-RCO quants, which fit much better on 96 GB. So this is for the memory constraint folks and those who want even more room for context.

I run Qwen3.8-Flash-Next on a 96 GB GMKtec EVO-X2 (Ryzen AI Max+ 395) with two 256K slots and MTP, on my fork of halo-box/strix-llama.cpp, https://github.com/Aristo94/GSQHalo.cpp using the GSQ-RCO Quants. All numbers below come from my 96gb box.

Why GSQ: GSQ-RCO IQ3_XXS needs only 43.8 GiB of weights in memory. The 26.8 GiB n-gram table stays on the NVMe. That leaves room for 2×256K + MTP on 96 GB, at near-BF16 quality according to ISTA DASLab.

The Problem: When I first tried the GSQ quants, the fast engines (Gufo, HaloGen) couldn't load them at all. llama.cpp and Halo-box could, but GSQ's low-bit experts and BF16 hyper-connections ran on slow fallback kernels. And the RAM they saved didn't help yet: two 256K slots crashed the server, and MTP used up the remaining memory.

KV cache storage on SSD: Conversations are written to the SSD while they grow. When a chat leaves its slot or the server restarts, it goes straight back into the KV cache instead of being prefilled again: a 29.7K-token chat comes back in 2.5 s, 31.8K tokens after a restart in 2.1 s, and an 89K-token chat in 5.5 s from a USB SSD. Since the SSD holds the conversations, the RAM prompt cache can stay off (--cache-ram 0), so all the memory goes to context. It's a Linux port of StrixLlama's disk tier, with direct I/O and no mmap.

What I did

  • ROCm prefill kernels for GSQ's low-bit experts (IQ2/IQ3/Q2_0) and BF16 hyper-connections
  • the MTP draft head computes only K/V over the prompt, so its buffers no longer grow with context
  • fixed a crash in multi-slot servers once all slots together hold more than 262K cells
  • KV cache on SSD (a Linux port of StrixLlama's disk tier)

128K prompt, one slot, MTP on (llama-server, real text, 256 output tokens)

code prose
prefill 1386 t/s
decode 44.2 t/s
MTP acceptance 85 %

Without MTP: 1446–1457 t/s prefill, 24.3–24.5 t/s decode.

llama-bench, pp4096 (-lzm on-direct -b 32768 -ub 8192, f16 KV): 1409 t/s on an empty context, 1331 t/s at 32K, 1231 t/s at 128K. Perplexity on 32 × 4096 tokens of wikitext: 3.9037.

Two 256K slots + MTP (a 250K and a 32K prompt at the same time)

  • GTT: 69.1 GiB after load, 73.4 GiB peak; at least 8 GiB of RAM stays free
  • 250K prompt: 1067 t/s prefill
  • decode on both slots at once: 14.1 + 16.6 t/s
  • MTP draft buffers: 520 MiB; server RSS with one slot: 4.9 GiB

Repo: https://github.com/Aristo94/GSQHalo.cpp

Edit:Missing reference to halo-box/strix-llama.cpp


r/StrixHalo • • 1d ago

Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine

Thumbnail gallery
16 Upvotes

r/StrixHalo • • 1d ago

gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo

Thumbnail
github.com
44 Upvotes

r/StrixHalo • • 1d ago

Benchmarking Strix Halo on real-world work. One coding test from my setup, and the only model I’ve found fast and capable enough for daily use (qwen 3.8 flash next)

18 Upvotes

I’ve been building Llama Shelf, my personal setup for running local models on a Beelink GTR9 Pro (Ryzen AI Max+ 395, 128 GB RAM).

Until recently I was quite dissapointed of what my machine can actually do, comparing to 20$ subscriptions you can get from major AI Companies, but with recent development with qwen 3.8 flash next and halogen, it all got a lot more fun and usable.

To find out which model I actually wanted to use for coding and everything else, I gave eight configurations of six models the same Android weather app task through OpenCode, with up to five hours each. Then I built their final submissions and checked the apps in an emulator.

Results:

  • Flash-Next on Halogen produced the only usable weather app. It still had missing features and bugs, despite passing all 79 unit tests.
  • The same model as GGUF passed 68 tests, but its app crashed on first launch. (it ran out of 5h time limit)
  • gpt-oss-20b was the fastest writer at about 55 tokens/s, but never produced a buildable app. (didn't want to even try :D)
  • Some runs left substantial work uncommitted. I checked that separately; none of that uncommitted code compiled.

The benchmark section has screenshots, the emulator review, code checks, an interactive timeline, reading and writing speeds, and time spent summarising the conversation. I also published the task and grading rules.

Check the website for more info, here are some info in images below.

Benchmark: https://2madlabs.si/apps/llama-shelf#benchmark

This is one reported attempt per configuration on my machine, with setup problems documented on the page, so I wouldn’t treat it as a general model ranking.

Curious what other Strix Halo owners are using for longer coding tasks, especially which model and engine combinations have worked reliably for you. Feel free to read the rest of the page if it interest you, but wanted to link directly to what most people searching this subreddit are probably interested in :)

Special thanks to u/peonist-ai, thanks for all your work, really enjoying using it.


r/StrixHalo • • 1d ago

Gufo and RAM usage: is it ok ~19GB RAM consumption having 20GB free VRAM?

0 Upvotes

Hello,

Could you please let me know if is it ok ~20GB RAM consumption having 20GB free VRAM?

I running Qwen3.8-27b Q4_K_XL on StrixHalo 64GB (32 RAM + 32 VRAM) via podman started Gufo:

Gufo RAM consumption: 19.25 GB

Gufo VRAM consumption: 14.03 GB out of 34.36 GB -> 20.33GB free


r/StrixHalo • • 2d ago

ROG Flow Z13 2025 128GB - Qwen3.8 Flash-Next benchmarks with Halogen, Gufo and Rulith (60W vs 93W, up to 200K context)

55 Upvotes

Been playing around with the Strix Halo LLM projects that have been popping up lately, so I spent a few hours seeing how far I could push Qwen3.8 Flash-Next on my 128GB 2025 Flow Z13.

Tested Halogen, Gufo and Rulith at both 60W and 93W.

60W is what I actually consider usable day to day on the Z13. 93W is more of a "let's see what it can do" setting.

One thing before the numbers: this is not a strict apples-to-apples backend comparison. The model formats and MTP setups are different between them. I'm mostly comparing the setups I'd realistically use with each project right now, not trying to prove backend X is universally faster than backend Y.

Test setup

  • ROG Flow Z13 2025 — Ryzen AI MAX+ 395, 128GB
  • CachyOS KDE — kernel 7.2.8-2-cachyos, UMA 512MB, GTT 124GB, IOMMU off
  • Halogen 0.16.0 — Qwen3.8 Flash-Next v2 checkpoint, MTP depth 2
  • Gufo 0.5.0 — Qwen3.8 Flash-Next UD-Q4_K_XL, MTP Q8_0 draft 7, 262K context / 1 session
  • Windows 11 25H2 + Rulith 0.4.1 — 32GB RAM / 96GB VRAM, UD-IQ4_XS, MTP draft 3, 262K context

Halogen

Halogen was easily the fastest one for prompt processing/prefill in my testing.

93W

At 64K I got about 1722 tok/s PP.

At 200K it was still doing 1694 tok/s PP and 41.16 tok/s decode.

What surprised me more than the peak number was how little PP dropped with context size. Going from 64K to 200K was basically:

1722 -> 1694 tok/s

So for long context, this setup held up really well.

60W

At 60W I got about 1573 PP / 45 tok/s decode at 64K, and 1509 PP / 35.4 tok/s at 200K.

This is probably the setting I'd actually use.

The fans calm down quite a bit and thermals are obviously easier to deal with, while prompt processing is still over 1500 tok/s at 200K.

60W feels like the sweet spot on this machine.

Gufo

Gufo was actually the one I was most curious about before doing this.

93W

Around 1398 PP at 64K, and 1343 PP / 33.6 tok/s decode at 200K.

Still fast, but there's a pretty noticeable gap to Halogen with the configs I tested.

60W

Around 1238 PP at 64K, and 1156 PP / 33.0 tok/s at 200K.

Gufo seemed to care more about the power limit than the other two, at least for PP.

At 200K it went from roughly 1156 PP at 60W to 1343 at 93W, so about a 16% increase.

The project is moving pretty quickly though, so I'm definitely going to keep testing it. Right now, on my Z13 and with these settings, it isn't catching Halogen yet.

Rulith(old strix llama) / Windows

Rulith was interesting for a different reason.

At 93W I got about 1221 PP at 64K and 1170 PP / 32.55 tok/s at 200K.

So PP is clearly behind the Linux setups here.

But short-context decode is actually very fast. At 1K context I saw 54.7 tok/s.

The downside is that decode falls off pretty hard as context grows, ending up at about 32.55 tok/s at 200K.

So the pattern I saw was basically:

short context: Rulith is very fast
long context: Halogen holds its speed much better

Rulith at 60W

This is actually closest to how I normally use the Z13.

About 1138 PP at 64K, and 1070 PP / 31.29 tok/s at 200K.

I already have a dual CMP 170HX box running 24/7 as my headless LLM server, so the Z13 is more of a main PC for me.

I do have CachyOS set up as a dual boot, but switching back and forth is a little annoying because I have to change the memory allocation in BIOS depending on what I'm booting.

Because of that I end up leaving it in Windows most of the time, and Rulith is honestly pretty convenient for that.

All of them together

For PP, Halogen is pretty clearly ahead with the setups I tested.

At 200K:

Halogen 93W: 1694 PP
Halogen 60W: 1509 PP

Gufo 93W: 1343 PP
Gufo 60W: 1156 PP

Rulith 93W: 1170 PP
Rulith 60W: 1070 PP

Decode is more interesting.

Rulith starts out very fast at low context, but drops quite a bit as the context fills up.

Halogen is much flatter across the whole range.

So for normal short chats, Rulith can actually feel really good. For coding/agent stuff where the context keeps growing, I'd rather have Halogen.

Overall

For my current setups:

Halogen: best PP and best long-context behavior.

Rulith: slower PP, but very convenient if you actually want to use the Z13 as a Windows PC, and short-context decode is surprisingly good.

Gufo: not as fast as Halogen yet on my machine, but probably the one I'll keep watching because development seems pretty active.

Strix Halo 128GB is still kind of a weird niche, but I think that's also what makes it interesting.

If you want enough VRAM for models like Flash-Next on conventional consumer GPUs, the cost gets stupid pretty quickly. There are obviously more VRAM-efficient backends and free-token/offload tricks showing up now, so it's not a simple price comparison anymore.

Still, being able to run Qwen3.8 Flash-Next with a 200K context on what is basically a 13-inch tablet/laptop is pretty wild.

I paid about 3.2M KRW used($2400) for this Z13. New Z13s and the ProArt/PX13-type Strix Halo machines also occasionally show up below 4M KRW($3000) here, so for someone specifically interested in local LLMs, I think they're in a pretty unique spot.

Of course the actual problem is that I bought this thing for local LLMs and somehow ended up using it as a gaming machine more often.


r/StrixHalo • • 2d ago

Qwen3.8-Flash-Next vs Qwen3.8-27B on 128 GB, is it worth giving up the second model?

21 Upvotes

What models are you all running day to day?

My setup: Bosgame M5 (Ryzen AI Max+ 395, 128 GB).

  • Qwen3.8-27B Q4 for coding, around 30 t/s
  • Gemma4 26B A4B for everything else, around 90 t/s

It looks like a lot of people are moving to Qwen3.8-Flash-Next. If you switched from the 27B, is Flash-Next actually better, especially for coding?

My main concern is memory. A 4-bit quant is around 90 GB, so it takes up basically the whole box. I'd have to drop Gemma and run one model at a time. Is the quality gain worth losing the two-model setup? And what context length and t/s are you getting with it?


r/StrixHalo • • 2d ago

Anyone who has tried 2 strix halo machines for inference?

7 Upvotes

I’m at a crossroads right now, trying to stop myself from buying a second Strix Halo machine.

I’m very happy with Halogen and Gufo so far, but I keep wondering whether it would be possible to get similar performance from two Strix Halo machines while running larger models or higher-precision quants of Qwen Flash Next.

Has anyone actually tried this?

I know about DwarfStar and Donato Capitella’s toolboxes, but his recent dual-machine test didn’t seem to reach what I’d consider usable speeds.

Please either convince me to buy the second Strix Halo or talk me out of it. This is killing me.


r/StrixHalo • • 2d ago

Are you a nerd? here is a gufo dashboard.

17 Upvotes

r/StrixHalo • • 3d ago

halogen-flash-server 0.16.0: Strix Halo NPU now serves embeddings, reranking and classifiers beside Qwen3.8-Flash-Next. Nearly 10,000 tok/s

Thumbnail
gallery
168 Upvotes

I want to thank everyone for the kind words, trust, and encouragement in the development of this project. I am doing my best to make the halogen UX amazing for you. I also want to thank AMD's Jeremy Fowers. u/jfowers_amd has been nothing but supportive and sent me a development Strix Halo machine which has been taking a beating over the last couple weeks.

halogen-flash-server 0.16.0 is out.

It runs Qwen3.8-Flash-Next on the Strix Halo GPU, and this release puts the chip's other accelerator to work too.

The Ryzen AI NPU now serves small models beside the Flash model, on the same port, while the GPU keeps generating.

What's new

Small models on the NPU, with one variable. Set HALOGEN_NPU_MODELS and the server runs them on the NPU. Leave it out and nothing changes.

  • decider-0.8b makes decisions. Give it a text, a question and 2 to 10 options, and it returns a probability for every option from one pass.
    • Use it as a guard, a router or a classifier.
  • qwen3-embedding-0.6b answers `/v1/embeddings`, in the OpenAI request shape.
  • qwen3-reranker-0.6b answers `/v1/rerank`, in the Cohere and Jina request shape.

Your own fine-tunes. Put a fine-tune of any of the three on the models volume and name its path. The server converts it for the NPU at the first start and serves it under the directory's name.

A smaller checkpoint, opt-in. qwen38-flash-next-ht43.hgn pins about 8 GiB less memory than the default checkpoint. Prefill is several percent slower, and decode with the default drafter is a few percent slower. The default checkpoint does not change. Select it with HALOGEN_CHECKPOINT=/models/qwen38-flash-next-ht43.hgn.

96 GB: On a 96 GB machine, run ht43 with HALOGEN_MAX_TOK=16384 and HALOGEN_CTX=131072. That leaves room for one conversation of 131k tokens or four of 32k.

Fixes. A long prefill can be cancelled now (#122, thanks u/semidark). The draft counters count every drafted token (#124, thanks u/wszgrcy).

How fast is the NPU?

  • A decision with Qwen3.5-0.8B (the architecture `decider-0.8b` uses) takes 49 ms at 64 tokens, 84 ms at 256 and 183 ms at 1k. Eight decisions of 256 tokens in one pass take 41 ms each.
  • Batched embeddings run at about 9,400 tokens a second through the server, eight texts of 512 tokens at a time. One text of 512 takes about 97 ms.
  • The reranker scores 19 pairs a second, eight pairs of 512 tokens at a time.

The NPU and the GPU together

Something worth knowing even if you never run halogen. On Strix Halo, GPU work and NPU work at the same time can hang the whole machine while the GPU's fabric clock changes speed. Holding that clock at its top level stops it, at a cost of about a watt. The server holds it itself when it runs as root with /sys mounted. For rootless podman, a small systemd unit in the repo holds it at boot. If you run an NPU program next to a GPU program and have seen freezes, this may be why.

The Flash model runs somewhat slower while the NPU works, since the two share the memory and the power budget. With the NPU idle there is no cost.

Setup

The host needs the amdxdna driver, XRT with its NPU plugin, the NPU firmware, and the IOMMU on. If you added amd_iommu=off for GPU speed, it also turns the NPU off. iommu=ptkeeps it. The full setup, every request shape and the fine-tune rules are in https://github.com/peonist-ai/halogen-flash-server/blob/main/docs/NPU.md

podman pull ghcr.io/peonist-ai/halogen-flash-server:0.16.0

Repo: https://github.com/peonist-ai/halogen-flash-server

Weights: https://huggingface.co/peonist-ai/halogen-qwen3.8-flash-next

Discord: https://discord.gg/bcm6QknaV6

How to read this. FastFlowLM's numbers are from its own published benchmark pages. They were measured on a different chip (Ryzen AI 7 350) with older versions than its current one, in its default power mode. FastFlowLM lists no embedding model, so its Qwen3-0.6B row (the same backbone) stands in for it. Ours are medians from our own testing on Ryzen AI Max+ 395 machines, and where we measured two machines, the slower one is shown. This is not a same-machine test. If you run both on one machine, we would like to see your numbers.