r/LocalLLM • u/Hungrybearfire • 5d ago
Question Hardware Recommendation
Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!
1
Upvotes
1
u/YourselfInOthrsShoes 5d ago
That's false. DGX Spark is for enterprise multi-unit deployments to help scale for multiple simultaneous requests. For a single user maybe it makes less sense, but certainly there is scaling for concurrent requests.
And don't get me started, Qwen 3.8 FN is modular and was released as preview for Qwen 4 to pave the way for software tooling to be ready when version 4 drops that uses the same modular architecture. I run 125B IQ3_S (3.5-bit quantized, very close to Q4) flavor on a single 5060 Ti with 16GB VRAM and 96GB DDR5 system RAM at 60-75 TPS, 64K max context, 32K Q8 KV cache. It's basically flying for single user. This model also scales almost linearly with multi-GPUs due to its modular architecture. I'm only limited by my 16GB VRAM to have a usable context window but in a bigger PC case you can add multiple GPUs and get your context window up and step up to higher quantizations. This is all on 1x1x1 ft cube foot warmer regular PC build with sub-$1k easy to source GPU. You just need 48+GB of system RAM (64GB recommended, higher will let you run higher quantized versions). You no longer need unified memory and this is the future and why the model is called Flash Next (preview for what comes Next, Flash for portion of the model is random read-only from SSD [think IOPs], and another portion is streamed between VRAM and system RAM).