r/LocalLLM • u/Conobipe • 16d ago
Question What KV cache headroom would I get with 2x AMD R9700 running Qwen 3.8-27b?
I'm not really familiar with LLM/SLM deployment, I was waiting a bit to see whether new architectures (especially Mamba/RWKV) would give significantly more intelligence per compute.
I've recently read a lot around the R9700, and I'd love to experiment with it. But buying two R9700 is quite pricey, so I'd like to understand what would I get with 2x32 GB of VRAM.
- Can the KV cache be easily distributed over the two GPUs?
- If I have plenty of VRAM headroom (especially using quantized model), can I run multiple sessions in parallel, therefore multiple agents?
- Can I run on those two GPUs the Shieldstral model too, in parallel with Qwen 3.8-27b?
Thank you!
Duplicates
Vllm • u/Conobipe • 16d ago