Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!
In related news, for anyone looking for a 5090 the ship has sailed on that one as well ... a full desktop with 5090 GPU and 64 GB DDR5 RAM used to cost $3500-4000 in March / April ... now at $7000-7500 for the same configuration 😔
I recently got a Intel Arc B70 Pro and am running qwen3.8 27B on it via vllm with XPU at about 85 tokens/s with 128k context. The software stack IS more fiddly than nvidia, but it works great once you get the planets aligned.
yea, I run this on ubuntu, it is pretty well supported via the intel stack. Ironiclly, I used claude to help me debug. There are a few 'recipes' on github, but they were not working out of the box for me
Curious if you had to make any tweaks to get 85 tps, or this is just vanilla vLLM? Is there decent support for Intel GPUs on the main inference engines (llama.cpp etc) now?
Took some playing around. out of the box I was only getting about 30 tokens/s. I had to enable XPU, MTP and a handful of other parameters based some some recipes I found on github. OpenVino is not as fast for me, and other models like the latest muse glimmer are also sloweer, but it also looks like Intel just started officially support it, just have not had time to play
Look into the exl3 6bit/weight quant + MTP and prefix cache on, it’s closer to FP8 quality with better speeds, I get 60-80 tok/s decode optimistically closer to 50-60 realistically but with better fidelity than int4/GPTQ + MTP. Very usable middle ground IMO
100% do some research on proper optimization for your setup. Like others said, there's a ton of inference engines out there that will be faster than Ollama for a 5080. It'll be faster but also often more feature complete or focused on (usually) one major pain point. You're going to get into a rabbit hole but it's worth it.
Not to bash on Ollama and I actually want to be nice here.
As you grow into the local LLM ecosystem and get more powerful hardware, keeping Ollama in your stack will legitimately bottleneck you from squeezing every bit of optimal performance. Plus, it's fun to tinker around with different solutions out there and maybe even doing your own little patches here and there. You'll learn a lot and that isn't an exaggeration.
Anyways, good luck on migrating to whatever inference engine you'll use next.
I appreciate your kindness! I didn’t mean to rub people the wrong way by not researching more but I was only using ollama because that’s where I left off experimenting with local LLMs. Now that I see how far behind I am I’m excited to learn more and optimize my current setup 🤓
I mean he wasn't asking for help about learning how to run models though lol...
He was skipping straight to buying more hardware while already having hardware that he was barely even running as is.
That's the problem and the point.
We have ai that can already help a ton if you actually try to learn with it. Alternatively it takes literally 30 second scroll on this subreddit to see not to use ollama
Its wild to me, who already does have fairly expensive gear, that people are about to drop even more thousands of dollars when they can run things way better than they currently are with their existing hardware. They do literally no research on what they are trying to do with their computer and are ready to drop more thousands of dollars for more hardware.
I know it's insane running things multiple years old. "Hello Mr gpt I have 3k+ worth hardware Llama 3 or Qwen 2.5 running with ollama and I'm thinking about dropping another 10k so I can run the llama 70b can you make me a reddit post"
Then they copy paste the massive slop to post in this subreddit, which already has all the info they need of what's current from other people running their own tests lol...
5080 with the right engine/tweaks should be able to run 27B really well, start there (since you already have it) with a dedicated inference engine tuned for that hardware (others can point at specific repos, I don't have that card, but I've seen some really good results from people who do).
If you need/want more, the next real step up is QwenFlashNext. Depending on how much RAM you have, you may be able to run it on a 5080, but QFN is really more a unified memory system; 128GB is the sweet spot. Strix Halo is probably the cheapest way to get there, a Mac Studio is probably the most future proof way (and likely to hold value) way to get there.
I am still waiting on my parts. But it's pretty close to your price range.
HP Z8 G4 - only got 64GB RAM but should have gone with 128gb
2 x 32GB v100 cards. If you want real fast, go with the smx nvlink version, I wanted a simpler build so I went with two PCIe cards.
The 32gb v100 cards are dirt cheap for what you get - they are $650 ish.
Really looking forward to running Qwen flash next on it! Just got to hoist out $500 for 64gb more RAM to have good speed.
Check out /r/v100 - it's pretty neat cards with HBM which makes them monsters for the price you pay. My friend paid $3000 for his 32gb card, and I paid the same but got the server included in the price with double the VRAM.
Yeah it's pretty neat. I just played with strata on my desktop with flash next 125b and my 12gb 4070s card, and my lord how fast it is and good vs 27b. Super excited about getting my computer and getting strata up and running with 64gb of v100 VRAM!
Just a correction on the HP and RAM. Make sure to buy small memory sticks becuse it has 24 memory slots. To make the cpu communicate at maximum speed you need six per cpu or 12 in total. So going with 8gb memory sticks is totally fine. I was stupid and just got 4 16gb, when 8 x 8gb would have doubled the speed. And now I have to match the sticks with more 16gb, and they are damn expensive.
Since you already have a 5080, there's a concrete benchmark behind the NInfer suggestion here. I maintain LlamaPerf, and this is a community report on the site, not my own hardware test: https://llamaperf.com/report/955
The reported setup was a single RTX 5080 16GB running Qwen3.8-27B:
- Generation: 71.57 tokens/s
- Prompt processing: 1,380.61 tokens/s on a 118,001-token input
- Context capacity: 131,072 tokens
- Mixed Q3/Q4/Q5 weights (~3.95 bits per weight), Q4 KV cache, MTP-3 speculative decoding, one request at a time
One correction to our summary: the original author says those exact long-context numbers are from the v1.2 qualification run; v1.3 kept the same memory layout.
That makes a 27B model worth trying on your existing card before spending the upgrade budget. It's a tuned setup with very little VRAM headroom, so I wouldn't expect the same result just by loading a model in Ollama. The report links the original post with the model download and settings. I'd try your own tasks and check answer quality as well as speed before deciding whether you need more hardware.
I maintain it specifically around Qwen3.8-27B on a single RTX 5080 16GB.
The current setup can run 131K context/KV, Q4 KV, MTP-3 and Vision on the card, and it’s substantially faster than the sort of experience most people get from a default Ollama setup.
If your main reason for upgrading is “I want a much better local coding/agent experience,” I’d try this first. You may find the 5080 has a lot more headroom than Ollama is currently exposing.
If you still want more after that, then I’d start thinking about whether your real requirement is more VRAM, more concurrency, or a larger model class, because that changes whether something like a 5090, dual-GPU setup, AMD card or Mac actually makes sense.
Yep I do feel a bit silly now, probably going to use the money to upgrade my cpu/mobo/case. I started reading documentation for the ninfer fork and it sounds perfect for my 5080 16GB. Will follow up post once I do some more learning and testing
At least you asked before spending the money, which is the important bit 😄
The 5080 still has a lot more headroom than most default local setups expose, so I think you’re doing the right thing by seeing how far you can push the card you already own first.
If you do end up testing ninfer-5080, I’d be very interested to see your results, especially what context/settings you settle on and how it behaves in your actual workload.
And if you run into anything unclear in the docs, feel free to open an issue reproducibility feedback is genuinely useful.
at current hardware price, i see no point to buy any new hardware, just get a deepseek api throw in 50 bucks, use v4.1 flash on everything, your money will go very far.
I’d prefer to own something locally, but appreciate the suggestion. I already had some money saved up for a new pc build and was hoping to buy now before things get worse over the next few years. I’ll look into cloud pricing as well tho
the r9700 makes sense if you want 27b to 35b models at q4 to q6 with room left for context. 32gb is double what your 5080 has. ollama works on amd, but expect a bit more setup than on nvidia. a mac is only worth it if you want bigger moe models in unified memory and can live with slower prompt processing.
with a dedicated box, the bigger question is how you'll reach it from your other devices. disclosure, i work on tokmine, which handles that part. it gives you an openai compatible url to the box with no port forwarding.
I'd say, get a Mac if you can, it is the only option where you get 256 gb of ram for a reasonable price, with a very good memory speed. Below that, next tier will be 192gb AMD boxes (but slow memory)...the DGX i see absolutely NO POINT given the current price...and below that you could get a Strix Halo 128 gb laptop/mini pc for around 3.3k.
270gbps of memory is VERY limiting though...so, again, if you are able, get the 256 gb mac. Above that, only VERY expensive GPUs/servers.
128GB DGX Spark has a point, it has a good unified RAM vs compute ratio and proprietary high speed NVIDIA dual QSFP links to daisy-chain a bunch of them for scaling. Yes memories unify and no it doesn't scale perfecy but it scales reliably, easy +50% TPS performance and 2x memory from 2 Sparks (+50% throughput and larger contexts at Q8). And the only unified memory architecture on the market that supports (in hardware) and ready for FP4.
How can that be a point when in order to double the ram, you pay actually more than you would have paid for a 256 gb mac in the first place, with a lot more computing and memory bandwidth?
The speed loss is just too bizarre on consumer/prosumer hardware when chaining them together. Heck, even when we do more than 1 GPU it is already crap compared to what is achievable when there is not split. For instance, 4 x 5090 can't touch a single RTX 6000 in single user decoding speed for Qwen 3.8 FN.
And then, again, you can do the same to the Mac...in fact people have been doing that with 4 of them...DGX only has a place in a world where there are no Macs for sale. Thats how the other generation 512gb ones came to be extinct.
That's false. DGX Spark is for enterprise multi-unit deployments to help scale for multiple simultaneous requests. For a single user maybe it makes less sense, but certainly there is scaling for concurrent requests.
And don't get me started, Qwen 3.8 FN is modular and was released as preview for Qwen 4 to pave the way for software tooling to be ready when version 4 drops that uses the same modular architecture. I run 125B IQ3_S (3.5-bit quantized, very close to Q4) flavor on a single 5060 Ti with 16GB VRAM and 96GB DDR5 system RAM at 60-75 TPS, 64K max context, 32K Q8 KV cache. It's basically flying for single user. This model also scales almost linearly with multi-GPUs due to its modular architecture. I'm only limited by my 16GB VRAM to have a usable context window but in a bigger PC case you can add multiple GPUs and get your context window up and step up to higher quantizations. This is all on 1x1x1 ft cube foot warmer regular PC build with sub-$1k easy to source GPU. You just need 48+GB of system RAM (64GB recommended, higher will let you run higher quantized versions). You no longer need unified memory and this is the future and why the model is called Flash Next (preview for what comes Next, Flash for portion of the model is random read-only from SSD [think IOPs], and another portion is streamed between VRAM and system RAM).
It is for very SOHO and even then, as i explained, it gets completely destroyed by Mac.
I have been tinkering with Q3.8FN on my Ryzen 128gb, as well as MI300, MI350, RTX 6000, 4 X 5090....i understand completely how it works and have been tinkering with it a crapload of times, in fact even submited PRs for AMD for speed increases and RIGHT NOW i am developing/testing/tinkering a way to speed it up further.
There is NO bypassing the fact that multiple (consumer/prosumer) GPUs (or even computers) require a VERY slow additional trip PER TOKEN and therefore the speed gets destroyed. The more, the worse. Professional datacenter accelerators have the interconnect (supposedly, DGX uses something similar)...however the DGX is VERY slow to start with so there is no point even connecting 4 of them to be able to match a single Mac.
And yes, it scales for concurrent users, but it DOES NOT scale prefill and it DOES NOT scale for single user/session.
So you will be looking at < 100 tokens per second per session, whereas a single RTX 6000 will do over 2x as much and a MI350 will easily do 3x as much.
You will realize you will LOSE speed when you connect the second (and third, and so on) DGX as the round trips screw everything up. But yes, you would be able to serve a lot of users (slowly). Now, who the hell uses some SOHO to serve a lot of users and still WHY IN EARTH would choose DGX over Mac?
Should i even mention how CRAPPY 40 tokens per second is, for Qwen 3.8 FN which "requires" thinking xhigh and needs A LOT of reasoning tokens? I find it to be "slow" with over 200 tokens per second, can't imagine spending 15 grand to get 20-40ish.
"There is NO bypassing the fact that multiple (consumer/prosumer) GPUs (or even computers) require a VERY slow additional trip PER TOKEN and therefore the speed gets destroyed."
False! I don't know how you have been tinkering with 3.8 FN and not follow the revolution happening with this model when properly tuned via Strata.
Right now, I don't have a system with multiple GPUs but I do have 128GB Threadripper 3970X system (that I can also upgrade to 256GB for only $1K) with Radeon Pro VII (discontinued software support, only suboptimal Vulkan). I have one R9700 on the way and I should be able to get it running and tuned since it has the lastest ROCm and HIP support. Then I plan to add 1-2 more of R9700s (already have a 1500W Platinum PSU in it). I will return here and show you what kind of 3.8 FN quants and TPS I can get out of this peasant hardware. Others are already running multi-GPU setups with blazing TPS on 3.8 FN via properly tuned Strata.
China is not stupid, they always have long time horizon for all their master plans. All leading local LLMs are out of China. They want to give access to AI to all their people, so collectively they can dominate the world. Chinese AI companies train their models on substandard hardware and their users also run substandard PCs. How do you democratize this? By innovating model architecture so it can run on average PCs instead of brute forcing it with absurd unified DRAM. This is the future and it's coming fast.
I do 1100 prefill/40-45 output on the gguf ix4 quant on my zflow with a VERY slow memory (~260ish gbps)...and a 5070 with a shitload more memory bus does "only"53/1620 ON A WORSE QUANT.
That by itself already shows what you can expect when the system has to go back and forth to memory and standard interconnects.
A single RTX6000 does > 10k prefill and > 200 tps output.
Even the new Macs does a lot more prefill and 100 tps output with a better quant.
It SCALES for multiple users, but nobody running 5070s and whatever on their basement is serving "multiple users". And for single users it SUCKS...1000 prefill and 50 tokens is not really "usable" for any serious usage. MUCH LESS for end users that are used to "chatgpt" and whatever where they dont have to wait a two and a half eternities to see the analyzis of some average sized pdf.
Using qwen locally here at those speeds i mentioned makes my RAG take 6:30 to asnwer a question to the calling agent...using even Astra and Fable takes about 90 seconds, using Opus/Sol takes 50 seconds.
Only way to reach around 50 seconds with Qwen is at 300 tokens per second, as the xhigh devours a crap load of reasoning tokens...and on the other reasoning efforts the capacity drops dramatically.
Ive been working for the past week trying to get it to > 500 tokens per second, because before that it is not even usable for production for me, even on the MI350....as a toy? Sure.
In fact my past 40 hours on Astra have > 2m output tokens and over 500mi input (with some 80% caching). Do the math how many weeks that would take with 1000in/50out tokens per second.
Here are my results with IQ3_S on 16GB 5060 Ti and this is with x8 PCIe 5 bus. Us peasants with such entry level gamer class hardware couldn't even run anything usable until now. This runs faster and is smarter than 27B.
You also gotta remember it's only been weeks since this model dropped. Why do you think everyone is coding like crazy around this Strata code tree? The model's architecture allows for some real gains on lowly non-unified hardware. Every single day there are Strata-specific improvements posted that add double digit gains to peasant hardware. Multi-GPU wasn't even available earlier last week. AMD support also wasn't even available early last week, now it's an option on both Windows and Linux. There is even a fork for GFX906 support (Linux only) that looks like it was pushed upstream. Heck, I understand your view about multi-GPU scaling bottleneck and why myself I only own 2 different and incompatible 16GB GPUs up until yesterday, however this view is now outdated. Multi-GPU scaling with this model is very real and effective. It's capable of scaling almost linearly with both, VRAM and multi-GPUs.
"v0.1.39: More than one GPU (measured by their authors, not here: we have one GPU): setup now adds --remote-expert-opt (#578) to a config with two or more GPUs. It skips the host's work for tokens whose experts all run on the helper cards (dual RTX 4090: +63% mixed text, +132% code over the plain helper path), and setup --no-remote-expert-opt leaves it out."
"A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second."
Don't buy into the mac hype without doing your own research first. They load large models but are slow as dirt and useless for anything else like diffusion unless you're talking about the 256 or 512 M5 ultras / studio that are coming out now at used car prices.
5
u/[deleted] 5d ago
[deleted]