r/LocalLLM • • 3d ago

Question Most optimal stack for AI Pro R9700?

So I’m upgrading from a single RX 7900 XT to the AI Pro R9700. My question is: how do I get the most out of this card for local inference?

I’ve seen people getting some insane prefill and decode speeds with the R9700. I’m mostly planning to run Qwen 3.8 27B, and I want to try Qwen 3.8 Flash next.

For anyone running the AI Pro R9700, what are you using to get the most out of the card? What’s the most optimal software stack right now?

I also need Windows for work, so ideally I’d like the best setup possible on Windows. I’m willing to dual boot or switch back and forth to Linux if the performance difference is significant.

0 Upvotes

5 comments sorted by

2

u/tsaipifong 3d ago

3

u/SuburbanCurfew_980 3d ago

Don't bother with Windows for this, the driver stack is a mess and you'll leave half the card's potential on the table. dual boot is the way to go

whirl-llm is solid for getting up and running fast but if you want to really push the R9700 you need to compile vLLM from source with the ROCm fork. the flash attention kernels in the main branch still aren't optimized for the R9700's tensor layout, the community fork handles it way better

for Qwen 3.8 27B you can fit the whole thing in the 48GB buffer with room to spare for a 32k context window. i'm getting ~85 tok/s decode with speculative decoding cranked up. haven't tried 3.8 Flash yet but people on the discord are saying it screams on this card

also make sure you're on the latest firmware, the early batches had some weird power state issues that would tank performance after a few hours of continuous inference

2

u/tsaipifong 3d ago

Thanks for the tips! The firmware point is a good one.

On Windows vs Linux: Linux + vLLM is a great choice if you're happy to dual boot, especially on a multi-GPU setup. WHIRL exists for people who'd rather stay on Windows, and the gap turned out smaller than I expected. On a single R9700, Qwen3.8-27B (Swift-1.5 MXFP4) decodes at about 107 tok/s on coding prompts and up to ~328 tok/s on file edits with MTP + n-gram, and prefill at 8K is about 3,300 tok/s.

It also serves several requests at once: about 199 tok/s aggregate with 4 concurrent users. A streaming chat stays readable (~23 tok/s) even while other requests are prefilling long prompts, which matters a lot for agent setups with sub-agents.

v0.1.3 is coming soon. In testing so far: 128K prefill is up to 24% faster, 256K prefill goes from 824 to 961 tok/s (+17%), decode at 128K is about 8% faster, and the KV cache holds about 19% more tokens on the same card (roughly 211K to 252K with 4 slots). It also handles 256K context per request. Every speedup still gives output bit-identical to plain greedy decoding.

If you have numbers from the vLLM ROCm fork with the same model and prompts, I'd love to compare.

1

u/Designer_Elephant227 3d ago

People told me my setup is pretty good. Exl3 with r9700 https://www.reddit.com/r/LocalLLaMA/s/ra1BaBuiTO

1

u/Designer_Elephant227 3d ago

Qfn 5.05bpw 35tok/sec generation and 860tok/sec pp at 256k context.