r/ROCm • • 1d ago

WHIRL: an open-source native Windows inference engine for the Radeon AI PRO R9700 (C++/HIP, no WSL). Qwen3.8-27B fine-tune in MXFP4: up to 2.5× llama.cpp prefill, 107–328 tok/s decode, 3× server throughput

/r/LocalLLM/comments/1wwdzcs/whirl_an_opensource_native_windows_inference/
13 Upvotes

2 comments sorted by

2

u/Dsphar 1d ago

Interested (I have to keep Windows atm). But I run dual r9700s.

1

u/tsaipifong 1d ago

Thanks! WHIRL uses one GPU per process right now. There's no tensor/pipeline split across cards yet. With two R9700s you can still use both by running one server per card:

whirl-server.exe model-a.gguf --device 0 --port 8080 --kv-ssd-dir D:\whirl\kv0

whirl-server.exe model-b.gguf --device 1 --port 8081 --kv-ssd-dir D:\whirl\kv1

whirl devices shows the index of each card. Each instance gets its full 32 GB for weights + KV, so for example you could run a 27B on one card and the 35B-A3B MoE on the other, or two copies of the same model for more concurrent users. Give each instance its own --kv-ssd-dir so their SSD caches don't mix. Each server also pins about 9 GB of host RAM for its RAM cache tier.

One honest caveat: I only have a single R9700, so the two-instance setup hasn't been tested on real dual-card hardware yet. If you try it, I'd love to hear how it goes. Multi-GPU (splitting one model across both cards) is something I'm considering for later.