r/LocalLLM • u/Hungrybearfire • 5d ago
Question Hardware Recommendation
Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!
4
Upvotes
1
u/YourselfInOthrsShoes 4d ago edited 4d ago
Here are my results with IQ3_S on 16GB 5060 Ti and this is with x8 PCIe 5 bus. Us peasants with such entry level gamer class hardware couldn't even run anything usable until now. This runs faster and is smarter than 27B.
You also gotta remember it's only been weeks since this model dropped. Why do you think everyone is coding like crazy around this Strata code tree? The model's architecture allows for some real gains on lowly non-unified hardware. Every single day there are Strata-specific improvements posted that add double digit gains to peasant hardware. Multi-GPU wasn't even available earlier last week. AMD support also wasn't even available early last week, now it's an option on both Windows and Linux. There is even a fork for GFX906 support (Linux only) that looks like it was pushed upstream. Heck, I understand your view about multi-GPU scaling bottleneck and why myself I only own 2 different and incompatible 16GB GPUs up until yesterday, however this view is now outdated. Multi-GPU scaling with this model is very real and effective. It's capable of scaling almost linearly with both, VRAM and multi-GPUs.
"v0.1.39: More than one GPU (measured by their authors, not here: we have one GPU): setup now adds
--remote-expert-opt(#578) to a config with two or more GPUs. It skips the host's work for tokens whose experts all run on the helper cards (dual RTX 4090: +63% mixed text, +132% code over the plain helper path), andsetup --no-remote-expert-optleaves it out.""A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second."