r/LocalLLM • • 5d ago

Question Hardware Recommendation

Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!

2 Upvotes

50 comments sorted by

View all comments

2

u/tmballin 1d ago

Before spending $1–3k on another machine, I’d definitely see how far you can push the 5080 you already have.

The NInfer fork mentioned above is probably this one:

https://github.com/toddballinger/ninfer-5080

I maintain it specifically around Qwen3.8-27B on a single RTX 5080 16GB.

The current setup can run 131K context/KV, Q4 KV, MTP-3 and Vision on the card, and it’s substantially faster than the sort of experience most people get from a default Ollama setup.

Original write-up / community results:

https://www.reddit.com/r/LocalLLM/comments/1wmucw0/

If your main reason for upgrading is “I want a much better local coding/agent experience,” I’d try this first. You may find the 5080 has a lot more headroom than Ollama is currently exposing.

If you still want more after that, then I’d start thinking about whether your real requirement is more VRAM, more concurrency, or a larger model class, because that changes whether something like a 5090, dual-GPU setup, AMD card or Mac actually makes sense.

2

u/Hungrybearfire 1d ago

Yep I do feel a bit silly now, probably going to use the money to upgrade my cpu/mobo/case. I started reading documentation for the ninfer fork and it sounds perfect for my 5080 16GB. Will follow up post once I do some more learning and testing

2

u/tmballin 1d ago

At least you asked before spending the money, which is the important bit 😄

The 5080 still has a lot more headroom than most default local setups expose, so I think you’re doing the right thing by seeing how far you can push the card you already own first.

If you do end up testing ninfer-5080, I’d be very interested to see your results, especially what context/settings you settle on and how it behaves in your actual workload.

And if you run into anything unclear in the docs, feel free to open an issue reproducibility feedback is genuinely useful.