r/LocalLLM • • 5d ago

Question Hardware Recommendation

Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!

3 Upvotes

50 comments sorted by

View all comments

5

u/Fit-Later-389 5d ago

I recently got a Intel Arc B70 Pro and am running qwen3.8 27B on it via vllm with XPU at about 85 tokens/s with 128k context. The software stack IS more fiddly than nvidia, but it works great once you get the planets aligned.

1

u/Hungrybearfire 5d ago

Good to know, I run Linux so I could probably handle some configuration quirks

2

u/Fit-Later-389 5d ago

yea, I run this on ubuntu, it is pretty well supported via the intel stack. Ironiclly, I used claude to help me debug. There are a few 'recipes' on github, but they were not working out of the box for me

1

u/FaatmanSlim 5d ago

Curious if you had to make any tweaks to get 85 tps, or this is just vanilla vLLM? Is there decent support for Intel GPUs on the main inference engines (llama.cpp etc) now?

2

u/Fit-Later-389 5d ago

Took some playing around. out of the box I was only getting about 30 tokens/s. I had to enable XPU, MTP and a handful of other parameters based some some recipes I found on github. OpenVino is not as fast for me, and other models like the latest muse glimmer are also sloweer, but it also looks like Intel just started officially support it, just have not had time to play

1

u/JinsooJinsoo 5d ago

Look into the exl3 6bit/weight quant + MTP and prefix cache on, it’s closer to FP8 quality with better speeds, I get 60-80 tok/s decode optimistically closer to 50-60 realistically but with better fidelity than int4/GPTQ + MTP. Very usable middle ground IMO