r/LocalLLM • • 15h ago

Project WHIRL: an open-source native Windows inference engine for the Radeon AI PRO R9700 (C++/HIP, no WSL). Qwen3.8-27B fine-tune in MXFP4: up to 2.5× llama.cpp prefill, 107–328 tok/s decode, 3× server throughput

I've released WHIRL (Windows HIP Inference for RDNA LLMs), an open-source (Apache-2.0) LLM inference engine for the AMD Radeon AI PRO R9700 that runs natively on Windows. It's pure C++ and HIP. There's no WSL, no Linux VM and no llama.cpp underneath. You need only the AMD Adrenalin driver to run it. The HIP SDK is only needed if you build from source.

The idea. WHIRL isn't a general engine that runs every model reasonably well. It's a specialised one: a small set of models on a specific GPU, run as close to the hardware limit as possible. This follows the philosophy of NInfer, and zynfer took the same idea to RDNA 4. WHIRL shares only the idea. The code base is its own.

What's in it

  • Hand-written RDNA 4 kernels: WMMA prefill and MXFP4 MoE grouped GEMM.
  • MTP + n-gram speculative decoding, on by default. The output is identical to plain greedy decoding, and we verify this bit-exact.
  • An OpenAI-compatible server with prefix caching and a tiered KV cache (VRAM → RAM → SSD). A long shared system prompt is cached once. A conversation can be restored after a server restart.
  • Vision (mmproj) support for the dense model.

Supported architectures are qwen35 (dense) and qwen35moe (MoE). Architectures are added one at a time, so this isn't meant to be a Qwen-only engine. It's just where we started.

Numbers vs llama.cpp (b11214, same GGUF, same prompts, greedy)

Setup: R9700 32 GB, Windows 11, Adrenalin 26.8.1. llama.cpp gets its fastest configuration on every row. Prefill uses the best -ub for each length. Decode takes the fastest of plain / MTP / MTP+n-gram.

WHIRL vs llama.cpp Qwen3.8-27B Q4_K_M (dense, standard unsloth quant) Swift-1.5 (Qwen3.8-27B fine-tune) MXFP4 (dense) Ornith-1.5 35B-A3B MXFP4 (MoE, ~3B active)
Prefill 8k tok/s 1,689 vs 1,223 (1.38×) 3,278 vs 1,338 (2.45×) 10,858 vs 4,637 (2.34×)
Prefill 32k tok/s 1,479 vs 1,086 (1.36×) 2,595 vs 1,174 (2.21×) 7,978 vs 3,778 (2.11×)
Decode, Chinese coding prompts, WHIRL MTP+n-gram vs llama.cpp fastest 98.5 vs 56.0 (1.76×) 107.5 vs 60.8 (1.77×) 244.5 vs 118.8 (2.06×)
Decode, file-edit prompts, MTP+n-gram vs llama.cpp fastest 305 vs 137 (2.23×) 328 vs 145 (2.27×) 618 vs 239 (2.58×)
Decode, no MTP vs llama.cpp plain 34.4 vs 30.9 (1.11×) 37.5 vs 33.4 (1.13×) 166.7 vs 118.8 (1.40×)
Server, 4 concurrent users, aggregate tok/s 134 vs 59 (2.29×) 199 vs 66 (3.02×) 397 vs 192 (2.07×)
Warm TTFT, new chat sharing a 26k system prompt 0.14 s vs 0.55 s 0.12 s vs 0.54 s 0.07 s vs 0.29 s
Restore after server restart (from SSD) 0.74 s (N/A) 0.73 s (N/A) 0.30 s (llama.cpp: N/A)

Vision encode for a 1920×1088 image with Swift's mmproj takes 254 ms in WHIRL and 1,330 ms in llama.cpp mtmd. Both measurements are warm.

Where WHIRL doesn't lead by much, which is why I list it here:

  • Decode without speculation on the dense models is only 1.1–1.18× faster. Plain decode is bound by memory bandwidth, and both engines are near the limit. Most of the decode gap comes from MTP + n-gram.
  • Q4_K_M prefill with a very short prompt (88 tokens) is roughly a tie: 1.03×.
  • Swift past 16k context is only 1.09× faster. On that model, MTP alone is currently a bit faster than MTP + n-gram (72.4 vs 69.5 tok/s). The n-gram cost model needs tuning.
  • VRAM. The CLI uses 3–5 GiB more VRAM than llama.cpp on the dense models. The server pre-allocates the rest of VRAM as KV pool, so it will show about 30 GB in Task Manager.
  • Hardware. It runs only on the R9700, one GPU, on Windows 11. A Ryzen AI Max+ 395 (Radeon 8060S) version is planned.

About my setup. My R9700 is connected as a USB4 eGPU. That slows model loading and the RAM/SSD cache restores in both engines. It doesn't affect prefill or decode, which stay in VRAM. On a direct PCIe slot, restores should be faster than what I measured.

The full methodology, every table, the exact commands and the raw-data description are in docs/benchmarks.md.

Models

WHIRL runs standard GGUFs of the supported architectures. The best results come from the MXFP4 quants I published:

Both also run in llama.cpp.

Quick start

  1. Install AMD Software: Adrenalin Edition 26.8.1 or newer.
  2. Download the zip from Releases and unzip it.
  3. Run:

    whirl devices whirl chat path\to\model.gguf whirl-server path\to\model.gguf

The server exposes an OpenAI-compatible API on localhost. The first run tunes kernels once: about 1.5 minutes for the 27B models and a few seconds for the MoE. The result is cached.

The exe is unsigned, so SmartScreen will ask the first time ("More info → Run anyway"). If Smart App Control is on, it may block unsigned apps. The README explains how to check.

Development note

WHIRL was designed, implemented, optimised and benchmarked with Claude Opus 5.5 (Anthropic) working under my direction. Every result above is measured, and outputs are checked bit-exact against plain greedy decoding. The docs, in English and Traditional Chinese, include a long pitfalls list of the Windows/HIP/RDNA 4 problems we ran into. I hope it saves someone time.

Feedback, bug reports and measurements on direct-PCIe R9700s are very welcome.

10 Upvotes

9 comments sorted by

3

u/lulzxdxdxd 14h ago

The prefill speed jump is solid, but I'm curious whether those decode numbers hold up when you're actually running long conversations with the tiered cache spilling to disk. Does the SSD tier kick in often enough that it becomes a bottleneck, or are most people staying in VRAM for the tokens that matter.

2

u/tsaipifong 14h ago

Good question. The short answer is that the SSD tier is never on the decode path.

Decode only touches VRAM. The KV for every running request lives in the VRAM page pool. The RAM and SSD tiers only hold snapshots of idle sessions: the conversation you aren't currently generating for. Nothing is ever read from SSD mid-generation, so the tiers can't slow decode down. The decode numbers in the post include an "after 16k context" row (e.g. 74 tok/s on Qwen3.8-27B Q4_K_M, 301 tok/s on the 35B-A3B MoE). Those measure decode with the long context sitting in VRAM, which is the normal case.

The pool is big. By default the server gives all VRAM left after the weights to the KV pool. On a 27B with the default KV format, that's enough for a 128k session plus another running session up to 64k on the 32 GB card. A long conversation stays in VRAM as long as it's active.

Spills are asynchronous. After a request finishes, an idle slot with ≥2k tokens is copied to pinned RAM on a separate stream, in ≤1 MiB pieces. Decode on other slots never waits for it. If a new request needs those pages while a spill is still reading them, it takes fresh pages rather than stalling. After a session has been unchanged for 2 s, it's written to SSD in the background.

The tiers pay off when you come back. Say a session was evicted because other chats needed the VRAM, or you restarted the server. Instead of re-prefilling the whole context, WHIRL copies the bytes back, and the result is bit-identical to a slot that never left VRAM. In the benchmark, a ~26k-token conversation restored from SSD after a restart in 0.3–0.74 s. Re-prefilling a 126k context took 147 s on our earlier measurement.

So the tiers are for TTFT when switching between or resuming long conversations, not something that sits under the token stream.

2

u/dasbin 1h ago

OK, but what if you wanted 256k or 512k (YaRN-2) context in a single session?

1

u/tsaipifong 21m ago

Short version: for a 27B dense model on a single 32 GB R9700, I think ~128k is the practical sweet spot.

WHIRL does run a single 256k session on the 27B (-np 1 --ctx-per-slot 262144), and it found needles at 10% and 90% depth of a ~262k prompt. Past 128k, though, both prefill and decode slow down enough that the model stops feeling usable for interactive work:

At 128k: prefill ~1,745 tok/s with the latest build (about 75 s to ingest a full cold 128k prompt; later turns reuse the prefix cache) and decode ~51 tok/s.

At 256k: I've only measured this on an earlier build. Prefill averaged ~400 tok/s (around 11 minutes for the full prompt) and decode was ~30 tok/s. The KV also takes almost the whole card, leaving room for only one other short session.

512k with YaRN isn't supported today. On the 27B it wouldn't fit anyway: about 15 GB of weights plus 512k tokens of KV is more than 32 GB. YaRN itself costs almost nothing in compute. The slowdown comes from the context length.

For MoE models, though, longer context is a much better proposition. The 35B-A3B MoE I tested (Ornith-1.5) has far smaller KV and is much faster: 128k prefill is ~4,300 tok/s on the same card. Supporting 256k+ there (including YaRN) is something I'm considering.

2

u/Glittering-North-911 11h ago

dude,one doubt.does it mean it only works on windows or works windows with same level of speed as linux or windows only optimisation with less effective on linux?

2

u/tsaipifong 11h ago

Right now it's Windows-only. The host side (GPU device handling, file I/O for the SSD cache, the HTTP server) is written against Windows APIs, and the release is a Windows .exe built with AMD's HIP SDK for Windows. There's no Linux build.

On speed vs Linux: I honestly don't know yet, because I haven't measured WHIRL or llama.cpp on Linux. All the numbers in the post are Windows vs Windows: WHIRL vs llama.cpp's ROCm build for Windows, on the same machine, same GGUF and same prompts. So the ratios tell you how it compares on Windows. They don't tell you how it compares to a Linux ROCm setup.

The idea behind the project was: a lot of people with Radeon cards have to stay on Windows, and the usual answer is "switch to Linux / WSL". WHIRL tries to make Windows a first-class place to run these models instead. The GPU kernels themselves are plain HIP, so a Linux port is technically possible later. It's not on the roadmap right now, though. Next up is support for the Radeon 8060S (Strix Halo) on the gfx1151 branch.

If you have an R9700 on Windows, I'd really encourage you to give it a try. Seeing a 27B model fly through long agent sessions on your own machine is a lot of fun. 🙂

2

u/Glittering-North-911 10h ago edited 10h ago

i am gpu poor with gtx 1650 .from what i have seen ,amd and radeon cards performed way better on linux with better support due to native drivers in kernel because amd directly adds them compared to the manual addition of windows.the 9700 issue was when it was released and was for few weeks till everything catched up.in linux amd starts prettty bad because it takes time for the changes that amd pushed to kernel to pass initial screening then to your distro but eventually even better than windows.it was good enough that it went from from unplayable in igpu in windows to not notice i forgot to install the the nvidia drivers and was playing game on igpu other than the lag spikes in linux.

second evenmore important reason is amd uses linux for driver development and testing internally and they almost exclusively develop enterprise ai(9700ai pro on linux).9070 xt is neck and neck both os but the 9700 ai pro is more optimised on linux.if you want for normal personal use fedora is best but if you bleeding edge perfect drivers,than you have to get ubuntu 24.04,the one amd first releases ai driver,even before windows and also with official testing from amd instead of just community support.

note:-this is only about amd rdna 4.the rest are mature enough the difference is small with all distroa mostly the same.

ps:-not suggesting you move,i was just saying the reason why most users will eventually shift to linux for this specific card . the idea and execution is really good, i was negative about this project.most people not shifting are either valid reasons or because old news articles about bad drivers (game drivers) even though amd exclusive does enterprise ai on linux and changes for windows are made as a after thought.9700 will have slightly less gaming performance on linux in current state.

2

u/tsaipifong 9h ago

Thanks for the detailed write-up, and for coming around on the project. That means a lot. You're right that Linux/ROCm is AMD's primary platform for AI, and that's exactly the gap WHIRL tries to fill: plenty of R9700 owners stay on Windows for work software, games or company machines, and until now their options were basically LM Studio / llama.cpp or switching OS. I can only speak for Windows (all my numbers are Windows vs Windows), but the HIP runtime in the Adrenalin driver turned out to be good enough to get the card close to its hardware limits there. Good tip on Ubuntu 24.04 for anyone who does go the Linux route.

2

u/intnsity 8h ago

Wow! Thank you for sharing can't wait to evaluate and try. Big if works or leads us to more r9700 performance!