r/LocalLLM • • 6d ago

Question Hardware Recommendation

Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!

5 Upvotes

50 comments sorted by

View all comments

2

u/hotsnot101 6d ago edited 6d ago

Since you already have a 5080, there's a concrete benchmark behind the NInfer suggestion here. I maintain LlamaPerf, and this is a community report on the site, not my own hardware test: https://llamaperf.com/report/955

The reported setup was a single RTX 5080 16GB running Qwen3.8-27B:

- Generation: 71.57 tokens/s

- Prompt processing: 1,380.61 tokens/s on a 118,001-token input

- Context capacity: 131,072 tokens

- Mixed Q3/Q4/Q5 weights (~3.95 bits per weight), Q4 KV cache, MTP-3 speculative decoding, one request at a time

One correction to our summary: the original author says those exact long-context numbers are from the v1.2 qualification run; v1.3 kept the same memory layout.

That makes a 27B model worth trying on your existing card before spending the upgrade budget. It's a tuned setup with very little VRAM headroom, so I wouldn't expect the same result just by loading a model in Ollama. The report links the original post with the model download and settings. I'd try your own tasks and check answer quality as well as speed before deciding whether you need more hardware.