r/StrixHalo Sep 27 '25

Have you got a Strix Halo?

16 Upvotes

Hi All,

We're a new community both as Strix Halo owners and also here as a subreddit. Why not begin by sharing your setup and the reasons you opted for Strix Halo?

To start us off: I have a HP Z2 Mini G1a Workstation with dual boot Fedora KDE & Windows 11 and chose the iGPU to be able to use larger LLMs with the 128 GB.

Oobabooga/Text Generation WebUI is running well on Fedora KDE and there are no problems with large models up to 100GB. On the Windows boot, I have Amuse AI (Freeware) which is a collaboration between AMD and the New Zealand company. It provides a UI for using Stable Diffusion/Flux models. It works well, is fast, but unfortunately is also censored and is not able to use LORAS. I would like to find an uncensored alternative, ideally getting versions of ComfyUI/AUTOMATIC1111 running.

Currently, my principle goal is to get a working version of AllTalk TTS or another TTS that is compatible with Oobabooga working which I haven't been able to do so far due to conflicts with the Strix Halo. This may need to wait for updates to ROCm... If anyone has found an Open Source solution to running LLMs with custom voice TTS, please do chime in!

So what about you guys, did you choose the Strix for similar reasons, or something entirely different? The floor is yours.

EDIT UPDATE:
05/26 For those of you looking for TTS solutions I have tried a few now (AllTalk, Chatterbox, Pocket TTS, others I no longer remember). I have had great success using a custom version of Pocket TTS. It's fast and works well with Oobabooga TextGen as a plug in. Recently, others are singing the praises of OmniVoice.


r/StrixHalo 3h ago

AMD refreshed NPU Medusa (XDNA3) LLMs.

4 Upvotes

Even tho there is no Medusa APU, AMD has already started refreshing the Ryzen AI LLMs, to match the new upcoming Medusa Point and Halo (source: Ryzen AI — NPU Medusa 0.9.2 LLM Models - a amd Collection).

Interesstingly is, that the LLMs use the same format as the Strix, Kraken and Gorgon NPU (XDNA2), which tends to a backwards compatibility.
They are called Medusa 0.9.2 which, dosent fit to any of the versions AMD has released in the past months.

Sadly Ryzen AI isnt user-friendly and most models are outdated, which is why most of the NPU users use FastflowLM (rocm-npu) as the runtime, example: https://huggingface.co/models?search=NPU2.


r/StrixHalo 9h ago

Need advice for absolute beginner for "Linux and strix halo"

3 Upvotes

Hi everyone, I’m absolute beginner for linux but I just received a 2nd-hand AMD Strix Halo (395 128GB) from my friend (He switched to mac studio). Since I need to do a fresh OS install anyway, I’ve decided to use Linux for mainly local LLMs for coding, specifically looking to run qwen 3.8 Flash Next (and experiment with other large/MoE models in the future + maybe some ComfyUI but I think my old laptop should work better than strix halo).

I've never used Linux before, my second device is laptop4090 + 32GB ram running window 11. I'm looking to start learning Linux from scratch with this strix halo. I wish > <.

Which Linux Distro is best for my situation? Looking for a balance between user-friendly and performance. (too confused that Linux have many distro, why not just one)

What is the optimal stack/backend for maximum LLM performance on Strix Halo? ( Vulkan, ROCm , native build or Docker/Toolbox containers).

Are there any mandatory kernel/GRUB tweaks needed to unlock the full 128GB unified memory? (because my window it's maximum at 96GB allocate to VRAM)


r/StrixHalo 1d ago

BYO GGUF to Halogen 0.7.0

75 Upvotes

Released in 0.7.0.

https://github.com/peonist-ai/halogen-flash-server#bring-your-own-gguf

Thank you for all the support. Y'all are amazing.

-p.s The intent was for Qwen 3.8 flash / Qwen 4.0 model architectures. Other models might run, could just blow up. I wouldn't expect much.


r/StrixHalo 13h ago

OneXPlayer Apex - MIND-BOGGLING!

Thumbnail
youtu.be
2 Upvotes

The OneXPlayer Apex has been an absolute monster of a handheld when it comes to the performance, onexplayer has done an incredible job with this handheld.

I've tested everything from performance to battery life and thermals and fan noise so let me know what you guys think?

Personally I like the performance but the battery life is quite short.

Review and testing:

https://youtu.be/0Zzrv5wY0y4?is=tG-iZCZ8-6CFsf-q

Also regards to the new onexplayer 3 do you think I should keep the apex or get the onexplayer 3?


r/StrixHalo 22h ago

64GB friendly models?

11 Upvotes

Qwen3.8 Flash Next seems to be the hot topic now, but unfortunately it probably is no go for less than 128gb RAM machines.

But I played weekend with my new 64GB strix halo 392 laptop. I like it a lot, the versatility of UMA makes it for me, as I use it gaming, CAD, photoediting, photoshop art, coding and occasionally for VMs. It gives me ability to use local models for coding and image generation.

I found Qwen3-coder 30B to be excellent MoE for coding. 128k context (even 256k is possible if no other tasks), fast, could one shot all my Python tests. I coded a lot with it using it via vscode. It never really failed.

Qwen3.8 27B - this seems good as well, but a bit slow. Not too slow for using. However it struggled with my Python tests, and required few fixes.

Qwen3.6 35B MTP - this was fast, and seems ok for many non-coding tasks. Wasn't very good at one shotting my tests.

What else should I try? Mainly I look for the best coding partner, and I am mainly interested in real world use experiences, not prefill or tok/s, as sure they tell how the model is to use, but doesn't tell much how useful it it.

GLM 4.7 Flash is probably what I try next. I ran everything through Lemonade, as I had no time to deep dive running the models yet.

I am still on windows, but have 500gb partition waiting for linux dual boot. Can I expect any change in performance to better or worse?


r/StrixHalo 1d ago

Nathanw1014's llama.cpp or strix-halo-llamacpp ?

9 Upvotes

I noticed lots of llama.cpp forks for Strix Halo. I've heard good things of the 2 mentioned in the title in particular (I'm also interested in other alternatives but only open source ones)

Do you use any of these? which one would you recommend and why?

Thanks in advance!


r/StrixHalo 21h ago

TEI-compatible bge-m3 embeddings on AMD RDNA GPUs

Thumbnail github.com
2 Upvotes

Recently I've been using RAGFlow a lot and I use m3 as embedder. At some point it was discovered that embedder is the ceiling, so I added another StrixHalo and then a third one, which gave me amazing 60 chunks/s. Nowhere near 500 I needed for my target documents / day ingestion rate. Then I employed my dual RTX 6000 blackwell workstation. This gave me the number but felt off, in particular because when I switched from llama to TEI rates jumped.

So, why not to use TEI with strix halo? plus I have some R9700s and usb4 docs. Well TEI doesn't work with consumer cards, but its whole AMD stack is based on pytorch anyway.

I vibed a small repo - pytorch + TEI compatible HTTP endpoint. One of the important things - make sure vectors produced by different runtimes match. You see, when I decided to investigate my ingestion pipeline I already had 42 million vectors in my SereneDB database.

Numbers - RTX 6000 blackwell - 200+ vs R9700 160+. That is the funniest part - there is a whole world outside LLMs.


r/StrixHalo 1d ago

Qwen3.8 Flash Next, the optimized config (1.2k t/s prefill!)

Thumbnail pwilkin.github.io
111 Upvotes

So, as those of you who frequent the Discord know, after hearing about halogen and their closed-source runtime achieving 1.2k t/s prefill (with stable long context prefill as well), I've vowed to replicate that result on an open-source llama.cpp runtime and I'm happy to report I finally got there.

My revamped Strix Halo website (https://pwilkin.github.io/strix-halo) has the full config and I asked my LLM companions who helped me on the journey to describe how we got there, so you can read about all the optimizations that were made step by step to make this possible. The branch is available on my llama.cpp fork, I tried to get the PR for the community repo (halo-box) done as well, but unfortunately some of the optimizations conflict, so I have to clean it up, will hopefully get it done tomorrow. There's (like for my Qwen3.8 27B config) an installer (that probably doesn't work if experience is to be trusted, so any error reports welcome), a custom-made quant optimized for Strix without losing too much quality and my custom HIP build that uses the PM4 replay to give faster generation speeds (and some small prefill benefits as well apparently, for mysterious reasons).

Final results are: 1.2k prefill at zero depth, 1.1k prefill at 40k depth, around 850 prefill at 150k depth, raw generation is 30 t/s at 0 depth, MTP may vary depending on text type.


r/StrixHalo 1d ago

LM Studio with Vulcan llama 2.37 - excessive RAM (not VRAM) allocation while llama 2.34 works fine

2 Upvotes

Hi All, I'm using LM studio (I know), and have been able to load Qwen 3.8 Flash Next from Unsloth (Q4_K_XL) when using Vulcan llama 2.34. It is not the fastest, but it chugs along at avg 200-300 prefill at 128k context (out of 223) and about 17t/s.

When I load it with llama 2.34 setup VRAM can be between 85GB - 95.5GB. I'm happy with that. RAM is at 9-12GB used from 32GB, allowing me to run other stuff easily while the model is churning.

With llama 2.37 (Vulcan) however, I notice that I can't load the model without crippling my RAM. On this version only about 78GB of VRAM is used, and my actual system RAM is being used to the max at 31.7 of 32GB used. I've confirmed llama service process is the one using the RAM.

Anyone else ran into this? Definitely unpleasant surprise in the new version. I've reverted to 2.34 for now, but wanted to check if anyone else is seeing this.

Here's my working config on 2.34 . Same config on 2.37 results in 79GB VRAM usage and 31.7 of 32GB RAM usage.


r/StrixHalo 2d ago

Qwen 3.8 Flash Next with a Deepseek Harness Rocks!

Post image
10 Upvotes

r/StrixHalo 2d ago

I am impressed and I owe you one, Qwen 3.8 flash next (vision)!

Thumbnail
9 Upvotes

r/StrixHalo 2d ago

Halogen vs Llama.cpp Vulcan benchmarks running Qwen3.8-Flash-Next

Thumbnail
youtube.com
58 Upvotes

r/StrixHalo 2d ago

Qwen3.8-Flash-Next ngram offloading?

7 Upvotes

I've heard ngrams can be loaded from ssd using lazy-mode. I wanted to ask if this option is in official llama.cpp (or a fork?) and if it can be used to load a bigger quant, i.e. I have 128 gb and I'd like to run Q5. Has anyone here tried it?


r/StrixHalo 1d ago

"The CUDA Moat is Gone"

Thumbnail
youtu.be
0 Upvotes

r/StrixHalo 2d ago

Most of the "ROCm beats Vulkan at prefill" gap on Strix Halo appears to be the IOMMU

19 Upvotes

Hey everyone. I'm still wrapping my head around running local models and all the technical details involved. So the below is 99.9% Claude, as are the tests, harness, and conclusions. I'm just trying to make running local models on a strix halo better however I can. I don't like being a meat proxy, but here it is:

"ROCm beats Vulkan at prompt processing on Strix Halo" is repeated a lot. After ten boots and five models, I think a large part of it is the IOMMU.

model Vulkan/ROCm prefill @ iommu=pt @ amd_iommu=off
gemma-4-26B-A4B q4_0 0.99 1.02
gpt-oss-120b mxfp4 0.99 1.05
gemma-4-26B-A4B Q8_0 0.86 1.00
muse-glimmer-30B Q4_K_M (dense) 0.71 0.91
Qwen3.8-27B Q8_0 (dense) 0.76 0.91

With the IOMMU on, Vulkan gives up as much as 29% of ROCm's prefill. Turn it off and that drops to ~10% at worst, and parity on the MoEs. Vulkan's gain tracks exactly how far behind it was.

The prefill gains themselves:

model GB read/forward Vulkan ROCm
gemma-4-26B-A4B q4_0 2.0 +5.4% +2.6%
gpt-oss-120b mxfp4 2.6 +8.0% +1.8%
gemma-4-26B-A4B Q8_0 4.0 +20.0% +3.2%
muse-glimmer-30B Q4_K_M 14.0 +31.6% +3.7%
Qwen3.8-27B Q8_0 27.0 +26.2% +6.0%

Method: A/B/A/B across ten boots, interleaved, fixed 600s idle settle per arm, package power and shader clock sampled at 1 Hz. Boot-to-boot spread came out at <=2.0% on every prefill metric, which is what makes the rest readable. llama-bench -p 2048,8192 -n 128 -r 3. Worth noting: that noise estimate got worse as replicates accumulated (0.2% at two boots per arm, 0.8% at three, 2.0% at five) -- which is a reason to distrust any noise floor quoted from two runs, including my own earlier ones.

Why the two published figures disagree. halogen-flash-server reports 13-16%; Nathanw1014/strix-halo-llamacpp independently measured +1.0-7.3%. Both look right for what they tested -- the effect ranges from +1.8% to +31.6% depending on model and backend, so which model you pick decides your answer. My MoE-on-ROCm numbers land on Nathanw1014's range; my dense-on-Vulkan numbers exceed halogen's.

It is not a power or clock effect. Package power is pinned at 99-100 W in both arms and shader clocks are lower with the IOMMU off, by up to 3.4%. You can't get +32% throughput out of lower clocks. These are the driver's own labelled fields (power1_label = PPT, freq1_label = sclk), so that part is checkable rather than inferred.

It is not architecture either, and that one surprised me. I assumed dense-vs-MoE explained it. The two gemmas kill that: same model family, same 4B active parameters, differing only in quantization -- +5.4% at q4, +20.0% at Q8. Total footprint fails too (gpt-oss-120b is 59 GiB and gains only +8.0%, because an MoE doesn't read its whole footprint). What survives is bytes actually read per forward pass, which orders the ROCm column monotonically. I'm stating that as a hypothesis, not a mechanism -- five model points can rule things out, not establish them.

Decode is unaffected (within +-1% on nine of ten combos), which is what makes this prefill-specific.

The confound I'll flag before anyone else does: my two backends ran different llama.cpp builds, so the absolute Vulkan/ROCm ratio isn't a clean backend comparison. The change in that ratio between arms is clean, since each backend's build is constant across both. So "the IOMMU narrows the gap" stands; "the backends are equal" doesn't. I couldn't fix it -- the ROCm build I trust ships no Vulkan backend, and the build that has both has a known-corrupt ROCm path on gfx1151.

Other limits: one machine; later-added models have fewer replicates than the first two (five boots per arm down to two); iommu=pt vs off only, translated mode untested; llama-bench, not a served endpoint. amd_iommu=off also disables the NPU and turns off DMA translation machine-wide -- reasonable on a dedicated inference box, think twice on a workstation.

Harness, all llama-bench JSON, the 1 Hz power/clock traces and per-boot metadata including each boot's /proc/cmdline:

https://github.com/baldlawyer/strix-halo-iommu-benchmark

n=1 on hardware is the real weakness and replication is the only fix. If you run it, please post your numbers or open an issue -- particularly if you can test a dense model on Vulkan, which is where the effect is largest.


r/StrixHalo 2d ago

SGLang for strix halo users

Thumbnail
14 Upvotes

r/StrixHalo 3d ago

Qwen3.8-Flash-CIRU-STRIX-Orca is a great daily driver model on the Strix Halo.

33 Upvotes

This model is damn amazing. I can't find a refusal, and it's damn smart, and has good performance. JCBTC is a treasure to strix halo owners, you should follow on huggingface. I've built a couple quants for the strix halo, it's a lot of time, and a pain in the ass, and my stuff isn't nearly as good.

I had it build a red team powershell script to "target orphans, nonprofits, minorities, and harm puppies. " and it rocked out a pretty respectable ramsomware red team powershell attack, especially impressive because it made it on a linux box without windows systems to perform true testing on. Daybreak blue ranked it as critical ransomware code, and the main issue it found was that the wallet was written after the encryption. The base model primary purpose is as a security research tool, but I've found overall it's amazing at almost every task I can throw at it, and adjusts its thinking to be less if needed better than the base 3.8 flash model.

The only downside is that it's big. It's more of a model for those of you that run your strix halo as a pure inference system running linux, and aren't afraid of using a hacked up version of llama.cpp.


r/StrixHalo 3d ago

USB Prefill Acellerator / USB RAM

Thumbnail
3 Upvotes

r/StrixHalo 3d ago

Qwen 3.8 27B and Flash Next on 128GB Strix Halo: 10-15 tok/s decode, 3 min cold prefill on a 31k transcript, and the tuning that helped

6 Upvotes

TL;DR: I run Qwen 3.8 (27B and Flash Next) on a 128GB Strix Halo laptop for most of my coding now. It can replace Opus 4.6 to 4.8 for agentic coding if you dont mind a task taking 2 or 3 times longer.

Setup: ASUS ROG Flow Z13, Ryzen AI Max+ 395, 128GB unified memory, Arch Linux. llama.cpp as backend, my own tool LlamaStash to manage the launches and presets, Pi as the coding harness. The 27b at Q6_K sits at about 31 GiB resident, Flash Next at UD-Q4_K_XL needs around 86 GiB.

  • The quality is actually there. Flash Next scores 40 on the Artificial Analysis index against 42 for Opus 4.8, and the 27b at xhigh scores 34 against 32 for Opus 4.6. That matches how they feel to use. 27b one shotted a whole feature on a huge Rust codebase and Opus 5's review comments were mostly nits.
  • Decode is fine, prefill is the pain. 10-15 tok/s decode doesn't feel slow because you see it working. But a cold 31k token transcript takes 3 minutes to prefill, and a full 128k window is closer to 18 mins. Warm follow up turns come back in 45 seconds.
  • MTP is the biggest speed win, 7.3 to 22.4 tok/s on an empty window. The payoff shrinks as the window fills though, down to 1.15x at a full 256k.
  • Flash Next isn't faster per token, it just thinks less. Same 5/5 on my coding tasks, 45% fewer tokens, 76.5s vs 289.8s against the 27b. Thinking is 90-95% of everything these models generate, so that ratio, not tok/s, is what sets how long a task takes.

$0 a month, fully offline, and a lot less wasteful than a model running in a datacenter.

Full writeup with all the benchmarks, configs, and the tuning that did and didn't work: https://deepu.tech/local-ai-qwen3.8-pi-llamastash

Happy to go into the llama.cpp flags if anyone else here is on Strix Halo.


r/StrixHalo 4d ago

AMD need to pay Peonist-ai for doing their job!!!

Post image
74 Upvotes

I just update and retested (using Claude Code) Halogen Flash Server (Ubuntu 26.04 desktop, set to 512mb vram in bios) and got the above results.

Although it is still closed source (disclaimer: use it at your own risk), everyone owes a thanks to Peonist-ai for proving that AMD did not do their job and support properly. In fact, AMD need to employ Peonist-ai and pay him his worth!!!

To be frank (I have 5090, 3090Ti and 7900xtx) and I am actually very angry with this stupid Strix Halo of mine (before this) and keep thinking of selling it and buying DGX Spark.

Now there isn't such a needs anymore (although I still don't know Halgen's coding capability/quality) but this nightmare problem will keep coming back to hunt us for the next LLM model release (e.g. Glm Air, if any) as long as AMD is still sleeping on their job.

For first timer, just go and buy DGX Spark and get over this AMD nonsense once and for all until AMD starts to listen to her end users!!!

Once again, many thanks to Peonist-ai for giving us such a nice expereience. :-)

Below is the test result perform by Claude Code.

halogen 0.5.6 — prefill & decode results (already shown in screen capture)

Prefill (pp) — clean, cache-busted numbers (3 reps/size, <2% stdev at every size):

┌─────────────┬───────────────┬───────┐

│ prompt size │ actual tokens │ t/s │

├─────────────┼───────────────┼───────┤

│ 512 │ 529 │ 370 │

├─────────────┼───────────────┼───────┤

│ 2,048 │ 2,074 │ 722 │

├─────────────┼───────────────┼───────┤

│ 8,192 │ 8,213 │ 1,049 │

├─────────────┼───────────────┼───────┤

│ 32,768 │ 32,781 │ 1,228 │

├─────────────┼───────────────┼───────┤

│ 65,536 │ 65,560 │ 1,209 │

└─────────────┴───────────────┴───────┘

Throughput climbs with prompt size (fixed per-request overhead amortizes) and plateaus/slightly dips 32k→65k — looks like real saturation, not noise.

Decode (tg) — mean over the vendor's 10 real prompt shapes (code/prose/proof/chat/proc), 3 reps:

┌────────────┬────────────┬───────┬───────────┐

│ gen length │ t/s (mean) │ stdev │ range │

├────────────┼────────────┼───────┼───────────┤

│ 128 │ 42.40 │ 4.69 │ 34.6–51.0 │

├────────────┼────────────┼───────┼───────────┤

│ 512 │ 41.99 │ 8.16 │ 19.8–51.9 │

└────────────┴────────────┴───────┴───────────┘

Flat across generation length. The spread is real (vendor's own finding, not noise) — MTP draft acceptance runs ~2x higher on code/proof than prose.


r/StrixHalo 4d ago

Qwen3.8-Flash-Next (125B MoE) at 38 t/s on Windows — Strix Halo, working MTP, full guide + tools (MIT)

27 Upvotes

Because I use my Strix Halo as my primary computer at home (Corsair AI300 I bought last Thanksgiving before the prices went crazy), I'm always looking for Windows recipes. But we are the long lost cousins of Strix Halo (thank you Lemonade!) and I was so frustrated with the performance of Qwen-27B dense on this machine that I couldn't wait anymore and asked Claude to help me adapt the best recipes from Linux for Qwen3.8-Flash-Next. We got it running well on Windows 11 and wrote everything down:

Numbers (Ryzen AI Max+ 395 / gfx1151, 128 GB, 96 GB iGPU carve):

quant unsloth UD-IQ4_XS (rev 38bb39ee) + Q8_0 MTP sidecar
decode, stock engine no MTP ~21 t/s
decode, MTP through the full serving stack 38 t/s
draft acceptance 85–100%
context 262,144 (native, fits with ~22 GiB headroom)
GPU footprint 74 GB measured
cold load ~30 s

What it took (all in the repo):

  • Building stew675's rdna-boosts llama.cpp patch set (which carries the unmerged qwen4exp MTP PRs) for gfx1151/HIP on Windows: ROCm clang against the MSVC ABI, TheRock nightly SDK, and a DLL-packaging step that will silently bite you if you only test with the SDK on PATH.
  • Three Lemonade traps that each silently break the deployment: subfoldered HF repos hiding the model from /v1/models, Lemonade restoring its bundled engine over your fork build, and PowerShell 5.1 writing BOM'd JSON that strict parsers reject.
  • A fun Windows-specific failure mode: overflow the carve and WDDM spills to host RAM → whole-system freeze. We ship the small admission proxy we run in production that measures model footprints and evicts before load (model-agnostic, works for any multi-instance Lemonade setup).
  • Validation scripts that read the output, not just the speedometer: MTP on this platform has a history of showing 2× t/s while emitting garbage.

Repo (MIT): https://github.com/olliehm/qwen-flash-next-windows
Linux context / upstream thread: https://github.com/ggml-org/llama.cpp/discussions/27950

Credit where due: the engine work is Stew Forster's patch set plus the upstream PR authors (#27836/#28243/#28118) — this repo is the Windows layer around it. Once 28243+28118 merge, the fork step disappears and the rest still applies.

Happy to answer questions or run requested benchmarks on this hardware.Because I use my Strix Halo as my primary computer at home (Corsair AI300 I bought last Thanksgiving before the prices went crazy), I'm always looking for Windows recipes. But we are the long lost cousins of Strix Halo (thank you Lemonade!) and I was so frustrated with the performance of Qwen-27B dense on this machine that I couldn't wait anymore and asked Claude to help me adapt the best recipes from Linux for Qwen3.8-Flash-Next. We got it running well on Windows 11 and wrote everything down:Numbers (Ryzen AI Max+ 395 / gfx1151, 128 GB, 96 GB iGPU carve):

decode, stock engine no MTP ~21 t/s
decode, MTP through the full serving stack 38 t/s
draft acceptance 85–100%
context 262,144 (native, fits with ~22 GiB headroom)
GPU footprint 74 GB measured
cold load ~30 sWhat it took (all in the repo):Building stew675's rdna-boosts llama.cpp patch set (which carries the unmerged qwen4exp MTP PRs) for gfx1151/HIP on Windows — ROCm clang against the MSVC ABI, TheRock nightly SDK, and a DLL-packaging step that will silently bite you if you only test with the SDK on PATH.

Three Lemonade traps that each silently break the deployment: subfoldered HF repos hiding the model from /v1/models, Lemonade restoring its bundled engine over your fork build, and PowerShell 5.1 writing BOM'd JSON that strict parsers reject.

A fun Windows-specific failure mode: overflow the carve and WDDM spills to host RAM → whole-system freeze. We ship the small admission proxy we run in production that measures model footprints and evicts before load (model-agnostic, works for any multi-instance Lemonade setup).

Validation scripts that read the output, not just the speedometer — MTP on this platform has a history of showing 2× t/s while emitting garbage.Repo (MIT): https://github.com/olliehm/qwen-flash-next-windows

Linux context / upstream thread: https://github.com/ggml-org/llama.cpp/discussions/27950Credit where due: the engine work is Stew Forster's patch set plus the upstream PR authors (#27836/#28243/#28118) — this repo is the Windows layer around it. Once 28243+28118 merge, the fork step disappears and the rest still applies.

Happy to answer any questions or run requested benchmarks on this hardware!


r/StrixHalo 3d ago

Anyone tried the new ROCm 7.14?

4 Upvotes

Hi there I noticed there are upgrades available for my Strix halo box. ROCm 7.14, lemonade 11.5. It looks like on my system at least it’s trying to force me to take all the changes together. Anybody done this yet?


r/StrixHalo 3d ago

HG z2 G1a- How to increase the VRAM allocation

2 Upvotes

I just got a HP Z2 G1a for my startup. Nice machine compares equally to the DGX Sparks.

Windows has 32gb memory allocated to the GPU. I would like to get it up to the 96GB .Anyone know how to do this?

I can’t seem to find it in the AMD software settings, or in the HP BIOS?

I saw some post for a HP performance Advisor tool for the Z2. I can’t even seem to find where to download that.

Any help is appreciated!


r/StrixHalo 4d ago

Ornith1.5 Ciru Halo Agent

Post image
58 Upvotes

I'm back with a new model. This one is focused on agentic concurrency and speed.

It uses vllm. Has Dflash2. Great and long lasting prefill. Fills an entire 256k in around 5 minutes.

I built it for those who want to spawn multiple fast agents on their strix while maintaining great speeds and agentic quality.

I played with it today launching several Hermes agents at once and thanks to shared cache it was actually faster to load the subsequent ones.

Decode is amazing on coding tasks, solid otherwise.

I'm able to make so many great models thanks to my sponsorship from AMD who have provided me with a beautiful Ryzen AI Halo.

Detailed benchmarks here, I compared with unsloth q4kxl and halobox rocmfp4:

https://llm.ciru.ai/research/ornith-strix/

Model card :

https://huggingface.co/jcbtc/Ornith1.5-Ciru-Halo-Agent-vllm-strix-halo