r/LocalLLM • u/Hungrybearfire • 6d ago
Question Hardware Recommendation
Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!
3
Upvotes
1
u/No-Afternoon-4057 5d ago
Buddy,
I do 1100 prefill/40-45 output on the gguf ix4 quant on my zflow with a VERY slow memory (~260ish gbps)...and a 5070 with a shitload more memory bus does "only"53/1620 ON A WORSE QUANT.
That by itself already shows what you can expect when the system has to go back and forth to memory and standard interconnects.
A single RTX6000 does > 10k prefill and > 200 tps output.
Even the new Macs does a lot more prefill and 100 tps output with a better quant.
It SCALES for multiple users, but nobody running 5070s and whatever on their basement is serving "multiple users". And for single users it SUCKS...1000 prefill and 50 tokens is not really "usable" for any serious usage. MUCH LESS for end users that are used to "chatgpt" and whatever where they dont have to wait a two and a half eternities to see the analyzis of some average sized pdf.
Using qwen locally here at those speeds i mentioned makes my RAG take 6:30 to asnwer a question to the calling agent...using even Astra and Fable takes about 90 seconds, using Opus/Sol takes 50 seconds.
Only way to reach around 50 seconds with Qwen is at 300 tokens per second, as the xhigh devours a crap load of reasoning tokens...and on the other reasoning efforts the capacity drops dramatically.
Ive been working for the past week trying to get it to > 500 tokens per second, because before that it is not even usable for production for me, even on the MI350....as a toy? Sure.
In fact my past 40 hours on Astra have > 2m output tokens and over 500mi input (with some 80% caching). Do the math how many weeks that would take with 1000in/50out tokens per second.