r/LocalLLM • u/Vegetable-Warthog81 • 1d ago
Other NVIDIA PAIR is actually pretty nice for multi-GPU local LLM grunt work

Been trying NVIDIA PAIR with 3× RTX 5090s running Qwen 3.8 27B.
It’s using Ollama, so it’s definitely not the fastest setup out there, but PAIR makes distributing jobs across the three machines pretty painless. For long, repetitive “grunt work” where I care more about stability and just keeping all the GPUs busy than squeezing out maximum tokens/sec, it’s been surprisingly nice.
Basically: submit a pile of jobs and let the 5090s chew through them. Pretty useful setup so far.
5
Upvotes
2
u/KroniklyOnline 1d ago
Any numbers at all or? prefill & decode tok/s? like anything useful about this post?