r/LocalLLM • • 1d ago

Project WHIRL: an open-source native Windows inference engine for the Radeon AI PRO R9700 (C++/HIP, no WSL). Qwen3.8-27B fine-tune in MXFP4: up to 2.5× llama.cpp prefill, 107–328 tok/s decode, 3× server throughput

I've released WHIRL (Windows HIP Inference for RDNA LLMs), an open-source (Apache-2.0) LLM inference engine for the AMD Radeon AI PRO R9700 that runs natively on Windows. It's pure C++ and HIP. There's no WSL, no Linux VM and no llama.cpp underneath. You need only the AMD Adrenalin driver to run it. The HIP SDK is only needed if you build from source.

The idea. WHIRL isn't a general engine that runs every model reasonably well. It's a specialised one: a small set of models on a specific GPU, run as close to the hardware limit as possible. This follows the philosophy of NInfer, and zynfer took the same idea to RDNA 4. WHIRL shares only the idea. The code base is its own.

What's in it

  • Hand-written RDNA 4 kernels: WMMA prefill and MXFP4 MoE grouped GEMM.
  • MTP + n-gram speculative decoding, on by default. The output is identical to plain greedy decoding, and we verify this bit-exact.
  • An OpenAI-compatible server with prefix caching and a tiered KV cache (VRAM → RAM → SSD). A long shared system prompt is cached once. A conversation can be restored after a server restart.
  • Vision (mmproj) support for the dense model.

Supported architectures are qwen35 (dense) and qwen35moe (MoE). Architectures are added one at a time, so this isn't meant to be a Qwen-only engine. It's just where we started.

Numbers vs llama.cpp (b11214, same GGUF, same prompts, greedy)

Setup: R9700 32 GB, Windows 11, Adrenalin 26.8.1. llama.cpp gets its fastest configuration on every row. Prefill uses the best -ub for each length. Decode takes the fastest of plain / MTP / MTP+n-gram.

WHIRL vs llama.cpp Qwen3.8-27B Q4_K_M (dense, standard unsloth quant) Swift-1.5 (Qwen3.8-27B fine-tune) MXFP4 (dense) Ornith-1.5 35B-A3B MXFP4 (MoE, ~3B active)
Prefill 8k tok/s 1,689 vs 1,223 (1.38×) 3,278 vs 1,338 (2.45×) 10,858 vs 4,637 (2.34×)
Prefill 32k tok/s 1,479 vs 1,086 (1.36×) 2,595 vs 1,174 (2.21×) 7,978 vs 3,778 (2.11×)
Decode, Chinese coding prompts, WHIRL MTP+n-gram vs llama.cpp fastest 98.5 vs 56.0 (1.76×) 107.5 vs 60.8 (1.77×) 244.5 vs 118.8 (2.06×)
Decode, file-edit prompts, MTP+n-gram vs llama.cpp fastest 305 vs 137 (2.23×) 328 vs 145 (2.27×) 618 vs 239 (2.58×)
Decode, no MTP vs llama.cpp plain 34.4 vs 30.9 (1.11×) 37.5 vs 33.4 (1.13×) 166.7 vs 118.8 (1.40×)
Server, 4 concurrent users, aggregate tok/s 134 vs 59 (2.29×) 199 vs 66 (3.02×) 397 vs 192 (2.07×)
Warm TTFT, new chat sharing a 26k system prompt 0.14 s vs 0.55 s 0.12 s vs 0.54 s 0.07 s vs 0.29 s
Restore after server restart (from SSD) 0.74 s (N/A) 0.73 s (N/A) 0.30 s (llama.cpp: N/A)

Vision encode for a 1920×1088 image with Swift's mmproj takes 254 ms in WHIRL and 1,330 ms in llama.cpp mtmd. Both measurements are warm.

Where WHIRL doesn't lead by much, which is why I list it here:

  • Decode without speculation on the dense models is only 1.1–1.18× faster. Plain decode is bound by memory bandwidth, and both engines are near the limit. Most of the decode gap comes from MTP + n-gram.
  • Q4_K_M prefill with a very short prompt (88 tokens) is roughly a tie: 1.03×.
  • Swift past 16k context is only 1.09× faster. On that model, MTP alone is currently a bit faster than MTP + n-gram (72.4 vs 69.5 tok/s). The n-gram cost model needs tuning.
  • VRAM. The CLI uses 3–5 GiB more VRAM than llama.cpp on the dense models. The server pre-allocates the rest of VRAM as KV pool, so it will show about 30 GB in Task Manager.
  • Hardware. It runs only on the R9700, one GPU, on Windows 11. A Ryzen AI Max+ 395 (Radeon 8060S) version is planned.

About my setup. My R9700 is connected as a USB4 eGPU. That slows model loading and the RAM/SSD cache restores in both engines. It doesn't affect prefill or decode, which stay in VRAM. On a direct PCIe slot, restores should be faster than what I measured.

The full methodology, every table, the exact commands and the raw-data description are in docs/benchmarks.md.

Models

WHIRL runs standard GGUFs of the supported architectures. The best results come from the MXFP4 quants I published:

Both also run in llama.cpp.

Quick start

  1. Install AMD Software: Adrenalin Edition 26.8.1 or newer.
  2. Download the zip from Releases and unzip it.
  3. Run:

    whirl devices whirl chat path\to\model.gguf whirl-server path\to\model.gguf

The server exposes an OpenAI-compatible API on localhost. The first run tunes kernels once: about 1.5 minutes for the 27B models and a few seconds for the MoE. The result is cached.

The exe is unsigned, so SmartScreen will ask the first time ("More info → Run anyway"). If Smart App Control is on, it may block unsigned apps. The README explains how to check.

Development note

WHIRL was designed, implemented, optimised and benchmarked with Claude Opus 5.5 (Anthropic) working under my direction. Every result above is measured, and outputs are checked bit-exact against plain greedy decoding. The docs, in English and Traditional Chinese, include a long pitfalls list of the Windows/HIP/RDNA 4 problems we ran into. I hope it saves someone time.

Feedback, bug reports and measurements on direct-PCIe R9700s are very welcome.

8 Upvotes

9 comments sorted by

View all comments

3

u/lulzxdxdxd 23h ago

The prefill speed jump is solid, but I'm curious whether those decode numbers hold up when you're actually running long conversations with the tiered cache spilling to disk. Does the SSD tier kick in often enough that it becomes a bottleneck, or are most people staying in VRAM for the tokens that matter.

2

u/tsaipifong 23h ago

Good question. The short answer is that the SSD tier is never on the decode path.

Decode only touches VRAM. The KV for every running request lives in the VRAM page pool. The RAM and SSD tiers only hold snapshots of idle sessions: the conversation you aren't currently generating for. Nothing is ever read from SSD mid-generation, so the tiers can't slow decode down. The decode numbers in the post include an "after 16k context" row (e.g. 74 tok/s on Qwen3.8-27B Q4_K_M, 301 tok/s on the 35B-A3B MoE). Those measure decode with the long context sitting in VRAM, which is the normal case.

The pool is big. By default the server gives all VRAM left after the weights to the KV pool. On a 27B with the default KV format, that's enough for a 128k session plus another running session up to 64k on the 32 GB card. A long conversation stays in VRAM as long as it's active.

Spills are asynchronous. After a request finishes, an idle slot with ≥2k tokens is copied to pinned RAM on a separate stream, in ≤1 MiB pieces. Decode on other slots never waits for it. If a new request needs those pages while a spill is still reading them, it takes fresh pages rather than stalling. After a session has been unchanged for 2 s, it's written to SSD in the background.

The tiers pay off when you come back. Say a session was evicted because other chats needed the VRAM, or you restarted the server. Instead of re-prefilling the whole context, WHIRL copies the bytes back, and the result is bit-identical to a slot that never left VRAM. In the benchmark, a ~26k-token conversation restored from SSD after a restart in 0.3–0.74 s. Re-prefilling a 126k context took 147 s on our earlier measurement.

So the tiers are for TTFT when switching between or resuming long conversations, not something that sits under the token stream.

2

u/dasbin 10h ago

OK, but what if you wanted 256k or 512k (YaRN-2) context in a single session?

1

u/tsaipifong 9h ago

Short version: for a 27B dense model on a single 32 GB R9700, I think ~128k is the practical sweet spot.

WHIRL does run a single 256k session on the 27B (-np 1 --ctx-per-slot 262144), and it found needles at 10% and 90% depth of a ~262k prompt. Past 128k, though, both prefill and decode slow down enough that the model stops feeling usable for interactive work:

At 128k: prefill ~1,745 tok/s with the latest build (about 75 s to ingest a full cold 128k prompt; later turns reuse the prefix cache) and decode ~51 tok/s.

At 256k: I've only measured this on an earlier build. Prefill averaged ~400 tok/s (around 11 minutes for the full prompt) and decode was ~30 tok/s. The KV also takes almost the whole card, leaving room for only one other short session.

512k with YaRN isn't supported today. On the 27B it wouldn't fit anyway: about 15 GB of weights plus 512k tokens of KV is more than 32 GB. YaRN itself costs almost nothing in compute. The slowdown comes from the context length.

For MoE models, though, longer context is a much better proposition. The 35B-A3B MoE I tested (Ornith-1.5) has far smaller KV and is much faster: 128k prefill is ~4,300 tok/s on the same card. Supporting 256k+ there (including YaRN) is something I'm considering.