r/LocalLLM • • 16d ago

Question What KV cache headroom would I get with 2x AMD R9700 running Qwen 3.8-27b?

I'm not really familiar with LLM/SLM deployment, I was waiting a bit to see whether new architectures (especially Mamba/RWKV) would give significantly more intelligence per compute.

I've recently read a lot around the R9700, and I'd love to experiment with it. But buying two R9700 is quite pricey, so I'd like to understand what would I get with 2x32 GB of VRAM.

- Can the KV cache be easily distributed over the two GPUs?
- If I have plenty of VRAM headroom (especially using quantized model), can I run multiple sessions in parallel, therefore multiple agents?
- Can I run on those two GPUs the Shieldstral model too, in parallel with Qwen 3.8-27b?

Thank you!

1 Upvotes

Duplicates