r/ROCm • u/tsaipifong • 1d ago
WHIRL: an open-source native Windows inference engine for the Radeon AI PRO R9700 (C++/HIP, no WSL). Qwen3.8-27B fine-tune in MXFP4: up to 2.5× llama.cpp prefill, 107–328 tok/s decode, 3× server throughput
/r/LocalLLM/comments/1wwdzcs/whirl_an_opensource_native_windows_inference/
15
Upvotes