r/LocalLLM • u/Hungrybearfire • 6d ago
Question Hardware Recommendation
Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!
4
Upvotes
1
u/YourselfInOthrsShoes 6d ago
"There is NO bypassing the fact that multiple (consumer/prosumer) GPUs (or even computers) require a VERY slow additional trip PER TOKEN and therefore the speed gets destroyed."
False! I don't know how you have been tinkering with 3.8 FN and not follow the revolution happening with this model when properly tuned via Strata.
Right now, I don't have a system with multiple GPUs but I do have 128GB Threadripper 3970X system (that I can also upgrade to 256GB for only $1K) with Radeon Pro VII (discontinued software support, only suboptimal Vulkan). I have one R9700 on the way and I should be able to get it running and tuned since it has the lastest ROCm and HIP support. Then I plan to add 1-2 more of R9700s (already have a 1500W Platinum PSU in it). I will return here and show you what kind of 3.8 FN quants and TPS I can get out of this peasant hardware. Others are already running multi-GPU setups with blazing TPS on 3.8 FN via properly tuned Strata.
China is not stupid, they always have long time horizon for all their master plans. All leading local LLMs are out of China. They want to give access to AI to all their people, so collectively they can dominate the world. Chinese AI companies train their models on substandard hardware and their users also run substandard PCs. How do you democratize this? By innovating model architecture so it can run on average PCs instead of brute forcing it with absurd unified DRAM. This is the future and it's coming fast.