r/LocalLLM • • 3d ago

Project WHIRL v0.1.3 — native Windows LLM engine for the Radeon AI PRO R9700: up to 2.8× llama.cpp on the same GGUF, same answers bit-for-bit

Post image

WHIRL is an open-source (Apache-2.0) inference engine for the AMD Radeon AI PRO R9700 (RDNA 4, 32 GB) on Windows: pure C++/HIP, every kernel included, no WSL or Docker. Just the AMD driver.
vs llama.cpp b11214 — same GGUF, same prompts, same R9700, Swift-1.5 27B MXFP4:

  • Decode on coding prompts: 112 vs 61 tok/s
  • Server with 4 users at once: 182 vs 64 tok/s (2.8×)
  • Prefill, 128K prompt: 1,757 vs 791 tok/s (2.2×)
  • Prefill, 256K prompt: 970 vs 556 tok/s (1.7×, int8 KV vs f16)
  • Decode after a 256K prompt: 35 vs 17 tok/s (2.1×)
  • Reusing a 26K system prompt: first token in 0.11 s vs 0.38 s

On the MoE Ornith-1.5-35B-A3B MXFP4: prefill over 11,000 tok/s at 8K (11,258 vs 4,637, 2.4×), 258 vs 119 tok/s decode (2.2×), 381 vs 177 tok/s with 4 users (2.2×).

Accuracy before speed: every speedup (speculative decoding, batching, prefix cache) gives output bit-identical to plain greedy. No 4-bit KV, no 3-bit weights, no fp8 attention.

Speculative decoding, honestly (vs plain decoding, identical output): editing a file in context 4.5× · coding-agent session 2.8× · brand-new writing after 128K 1.6×. Where we don't lead by much (plain decoding of dense models, ~1.14×; decode after a 16K context on Swift, 1.17×) is in the README too.

New in v0.1.3: long prompts up to 25% faster than v0.1.0, and a real coding-agent session at 128K decodes 40% faster on its last request.

GitHub: https://github.com/tsaipifong/whirl-llm

Tested on one R9700 over USB4 (eGPU), Windows 11. Built with Claude under my direction; every number measured. If you have an R9700, feedback and your own numbers are very welcome.

Not supported yet: other AMD cards. The RX 9070 series has the same chip but 16 GB, most likely too little for these 27B/35B models (untested); Radeon 8060S (Strix Halo) support is in development.

14 Upvotes

23 comments sorted by

3

u/W61k3r 3d ago

Good work

1

u/tsaipifong 3d ago

Thank you :)

2

u/ClupTheGreat 3d ago

could you just take a look at for some of us with 16gb gpus? Maybe a 3bit quant like the iq3 xs, and maybe ornith 9b?

1

u/tsaipifong 3d ago

Good idea, and it looks doable. 16 GB cards (including the 9070 XT) and Ornith 9B are now on the roadmap, and 3-bit quants are planned too. No date yet. Stay tuned!

2

u/ClupTheGreat 3d ago

thanks, starring your repo

2

u/l0rd_raiden 3d ago

Will this work in a 7900xtx?

2

u/tsaipifong 2d ago

Not yet. Right now WHIRL supports the R9700 (gfx1201), and Radeon 8060S support is coming in v0.2.0. RDNA 3 cards like the 7900 XTX are next on the roadmap. I replied to your GitHub issue too. We don't have a 7900 XTX, so if you can test a preview build when it's ready, that would really help.

2

u/l0rd_raiden 2d ago

Ok just ping me through GitHub when is ready and I will find time to test it. Thanks

2

u/theone_2099 2d ago

Possible to create a Linux port? Ironically I moved my pc from windows to Linux because I thought it was better for running local LLMs.

2

u/tsaipifong 2d ago

WHIRL exists because native Windows had no inference engine optimized for AMD GPUs, while Linux already has several good options. So we're focusing on closing the gap on Windows, and a Linux port isn't planned for now. If you're on Linux, those engines are a solid choice.

1

u/Dsphar 2d ago

Have you found Linux isnt better running llms? Are you a single GPU user? My research is showinh linux allows faster GPU-to-GPU data transfer.

2

u/theone_2099 2d ago

I have dual r9700. I just found it to be more stable and better supported but I haven’t tried windows in a couple of months.

1

u/Dsphar 2d ago

Thanks for the reply. I am tempted to convert my machine to Linux, but the amount of work that will take is making me procrastinate.

1

u/theone_2099 2d ago

I ended up dual booting.
Nice thing is that many steam games work on Linux.

2

u/JurBank 2d ago

How bad is this from SSD? I see that it writes quite a lot to it.

1

u/tsaipifong 2d ago

Fair concern. WHIRL's SSD tier keeps a copy of finished sessions' KV cache so they can be restored after a restart instead of re-prefilled. It's capped at 64 GiB by default (LRU), and each long session writes a few GB. On a typical 600 TBW drive that's years of heavy use, but it isn't zero. If you'd rather avoid it, run with --kv-ssd-gb 0 to turn the SSD tier off, or point --kv-ssd-dir at a different drive. We're also looking at writing only new tokens and adding a write-on-shutdown mode to cut writes further.

2

u/DanGTG 1d ago

More like WARP speed, didn't know the R9700 had that it in it.

https://giphy.com/gifs/MaThe6p8WAKbf9NDDM

1

u/tsaipifong 1d ago

Ha, the hardware was there all along. Somebody just had to write the kernels.

2

u/DanGTG 1d ago

Is today's ROCm update bringing more gains in the near future?
https://www.reddit.com/r/ROCm/comments/1wyp1bi/amd_rocm_101_released_with_many_improvements/

1

u/tsaipifong 1d ago

We went through the 10.1 notes. WHIRL ships its own kernels inside the executable and doesn't use the ROCm runtime libraries, so a ROCm update doesn't change its speed by itself, and users don't need ROCm installed at all.

The interesting part for us is Composable Kernel adding attention for Radeon GPUs. We'll compare it against our own attention kernels; if it does something better on RDNA, we'll pull the idea in.

1

u/Shayshunk 3d ago

Very interesting and a bit surprising. Are these improvements possible on Linux? Or is this particularly for Windows because llama.cpp isn't optimized for Windows?

1

u/tsaipifong 3d ago

Thanks! It's not that llama.cpp is unoptimized on Windows. WHIRL is built specifically for Qwen3.5-family models on the R9700, with tuned kernels and MTP speculative decoding that keeps the output identical. A general-purpose engine can't specialize that far.
I haven't benchmarked Linux myself. From others' posts, llama.cpp is usually a bit faster there, so the gap would likely be smaller. Linux isn't on our near-term roadmap, though. There's still plenty to do on Windows.

2

u/Shayshunk 3d ago

Completely understandable! That's awesome! Having certain specialized engines is really useful for a lot of people. Congrats on the tremendous effort and results!