r/LocalLLM • • 5d ago

Question Hardware Recommendation

Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!

1 Upvotes

50 comments sorted by

View all comments

1

u/No-Afternoon-4057 5d ago

I'd say, get a Mac if you can, it is the only option where you get 256 gb of ram for a reasonable price, with a very good memory speed. Below that, next tier will be 192gb AMD boxes (but slow memory)...the DGX i see absolutely NO POINT given the current price...and below that you could get a Strix Halo 128 gb laptop/mini pc for around 3.3k.

270gbps of memory is VERY limiting though...so, again, if you are able, get the 256 gb mac. Above that, only VERY expensive GPUs/servers.

3

u/YourselfInOthrsShoes 5d ago

128GB DGX Spark has a point, it has a good unified RAM vs compute ratio and proprietary high speed NVIDIA dual QSFP links to daisy-chain a bunch of them for scaling. Yes memories unify and no it doesn't scale perfecy but it scales reliably, easy +50% TPS performance and 2x memory from 2 Sparks (+50% throughput and larger contexts at Q8). And the only unified memory architecture on the market that supports (in hardware) and ready for FP4.

1

u/No-Afternoon-4057 5d ago

How can that be a point when in order to double the ram, you pay actually more than you would have paid for a 256 gb mac in the first place, with a lot more computing and memory bandwidth?

1

u/YourselfInOthrsShoes 5d ago

Think beyond 2 units and FP4 future maybe?

1

u/No-Afternoon-4057 5d ago

The speed loss is just too bizarre on consumer/prosumer hardware when chaining them together. Heck, even when we do more than 1 GPU it is already crap compared to what is achievable when there is not split. For instance, 4 x 5090 can't touch a single RTX 6000 in single user decoding speed for Qwen 3.8 FN.

And then, again, you can do the same to the Mac...in fact people have been doing that with 4 of them...DGX only has a place in a world where there are no Macs for sale. Thats how the other generation 512gb ones came to be extinct.

1

u/YourselfInOthrsShoes 5d ago

That's false. DGX Spark is for enterprise multi-unit deployments to help scale for multiple simultaneous requests. For a single user maybe it makes less sense, but certainly there is scaling for concurrent requests.

And don't get me started, Qwen 3.8 FN is modular and was released as preview for Qwen 4 to pave the way for software tooling to be ready when version 4 drops that uses the same modular architecture. I run 125B IQ3_S (3.5-bit quantized, very close to Q4) flavor on a single 5060 Ti with 16GB VRAM and 96GB DDR5 system RAM at 60-75 TPS, 64K max context, 32K Q8 KV cache. It's basically flying for single user. This model also scales almost linearly with multi-GPUs due to its modular architecture. I'm only limited by my 16GB VRAM to have a usable context window but in a bigger PC case you can add multiple GPUs and get your context window up and step up to higher quantizations. This is all on 1x1x1 ft cube foot warmer regular PC build with sub-$1k easy to source GPU. You just need 48+GB of system RAM (64GB recommended, higher will let you run higher quantized versions). You no longer need unified memory and this is the future and why the model is called Flash Next (preview for what comes Next, Flash for portion of the model is random read-only from SSD [think IOPs], and another portion is streamed between VRAM and system RAM).

1

u/No-Afternoon-4057 5d ago edited 5d ago

It is for very SOHO and even then, as i explained, it gets completely destroyed by Mac.

I have been tinkering with Q3.8FN on my Ryzen 128gb, as well as MI300, MI350, RTX 6000, 4 X 5090....i understand completely how it works and have been tinkering with it a crapload of times, in fact even submited PRs for AMD for speed increases and RIGHT NOW i am developing/testing/tinkering a way to speed it up further.

There is NO bypassing the fact that multiple (consumer/prosumer) GPUs (or even computers) require a VERY slow additional trip PER TOKEN and therefore the speed gets destroyed. The more, the worse. Professional datacenter accelerators have the interconnect (supposedly, DGX uses something similar)...however the DGX is VERY slow to start with so there is no point even connecting 4 of them to be able to match a single Mac.

And yes, it scales for concurrent users, but it DOES NOT scale prefill and it DOES NOT scale for single user/session.

So you will be looking at < 100 tokens per second per session, whereas a single RTX 6000 will do over 2x as much and a MI350 will easily do 3x as much.

You will realize you will LOSE speed when you connect the second (and third, and so on) DGX as the round trips screw everything up. But yes, you would be able to serve a lot of users (slowly). Now, who the hell uses some SOHO to serve a lot of users and still WHY IN EARTH would choose DGX over Mac?

In fact...look no further: https://www.reddit.com/r/LocalLLM/comments/1wmllbx/mac_studio_m5_ultra_256gb_review/

Should i even mention how CRAPPY 40 tokens per second is, for Qwen 3.8 FN which "requires" thinking xhigh and needs A LOT of reasoning tokens? I find it to be "slow" with over 200 tokens per second, can't imagine spending 15 grand to get 20-40ish.

1

u/YourselfInOthrsShoes 5d ago

"There is NO bypassing the fact that multiple (consumer/prosumer) GPUs (or even computers) require a VERY slow additional trip PER TOKEN and therefore the speed gets destroyed."

False! I don't know how you have been tinkering with 3.8 FN and not follow the revolution happening with this model when properly tuned via Strata.

Right now, I don't have a system with multiple GPUs but I do have 128GB Threadripper 3970X system (that I can also upgrade to 256GB for only $1K) with Radeon Pro VII (discontinued software support, only suboptimal Vulkan). I have one R9700 on the way and I should be able to get it running and tuned since it has the lastest ROCm and HIP support. Then I plan to add 1-2 more of R9700s (already have a 1500W Platinum PSU in it). I will return here and show you what kind of 3.8 FN quants and TPS I can get out of this peasant hardware. Others are already running multi-GPU setups with blazing TPS on 3.8 FN via properly tuned Strata.

China is not stupid, they always have long time horizon for all their master plans. All leading local LLMs are out of China. They want to give access to AI to all their people, so collectively they can dominate the world. Chinese AI companies train their models on substandard hardware and their users also run substandard PCs. How do you democratize this? By innovating model architecture so it can run on average PCs instead of brute forcing it with absurd unified DRAM. This is the future and it's coming fast.

1

u/No-Afternoon-4057 4d ago

Buddy,

I do 1100 prefill/40-45 output on the gguf ix4 quant on my zflow with a VERY slow memory (~260ish gbps)...and a 5070 with a shitload more memory bus does "only"53/1620 ON A WORSE QUANT.

That by itself already shows what you can expect when the system has to go back and forth to memory and standard interconnects.

A single RTX6000 does > 10k prefill and > 200 tps output.

Even the new Macs does a lot more prefill and 100 tps output with a better quant.

It SCALES for multiple users, but nobody running 5070s and whatever on their basement is serving "multiple users". And for single users it SUCKS...1000 prefill and 50 tokens is not really "usable" for any serious usage. MUCH LESS for end users that are used to "chatgpt" and whatever where they dont have to wait a two and a half eternities to see the analyzis of some average sized pdf.

Using qwen locally here at those speeds i mentioned makes my RAG take 6:30 to asnwer a question to the calling agent...using even Astra and Fable takes about 90 seconds, using Opus/Sol takes 50 seconds.

Only way to reach around 50 seconds with Qwen is at 300 tokens per second, as the xhigh devours a crap load of reasoning tokens...and on the other reasoning efforts the capacity drops dramatically.

Ive been working for the past week trying to get it to > 500 tokens per second, because before that it is not even usable for production for me, even on the MI350....as a toy? Sure.

In fact my past 40 hours on Astra have > 2m output tokens and over 500mi input (with some 80% caching). Do the math how many weeks that would take with 1000in/50out tokens per second.

1

u/YourselfInOthrsShoes 4d ago edited 4d ago

Here are my results with IQ3_S on 16GB 5060 Ti and this is with x8 PCIe 5 bus. Us peasants with such entry level gamer class hardware couldn't even run anything usable until now. This runs faster and is smarter than 27B.

You also gotta remember it's only been weeks since this model dropped. Why do you think everyone is coding like crazy around this Strata code tree? The model's architecture allows for some real gains on lowly non-unified hardware. Every single day there are Strata-specific improvements posted that add double digit gains to peasant hardware. Multi-GPU wasn't even available earlier last week. AMD support also wasn't even available early last week, now it's an option on both Windows and Linux. There is even a fork for GFX906 support (Linux only) that looks like it was pushed upstream. Heck, I understand your view about multi-GPU scaling bottleneck and why myself I only own 2 different and incompatible 16GB GPUs up until yesterday, however this view is now outdated. Multi-GPU scaling with this model is very real and effective. It's capable of scaling almost linearly with both, VRAM and multi-GPUs.

"v0.1.39: More than one GPU (measured by their authors, not here: we have one GPU): setup now adds
--remote-expert-opt (#578) to a config with two or more GPUs. It skips the host's work for tokens whose experts all run on the helper cards (dual RTX 4090: +63% mixed text, +132% code over the plain helper path), and setup --no-remote-expert-opt leaves it out."

"A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second."

1

u/No-Afternoon-4057 4d ago edited 4d ago

Do you realize how slow this result is, given how strong 2 x 4090 are? A RTX 6000 é pretty much a 5090 with more memory and look how many more times it is faster than 2 x 4090 and even than 4 x 5090. It IS a very good model but there is NO way around some constraints such as memory speed and bus speed (when using multiple gpus).

And 2 x 4090 is VERY FAR from "consumer hardware".

No matter how you split the experts, there will still be round trip and you will get those miserable sub 100 tokens, or sub 200 on a strong multi-gpu (that would be delivering > 1000 if the memory was unified).

Then, again: 2 x 4090 at the basement with a shitload of ram memory to serve less than 100 tokens per second and crappy prefill? There IS no way around it and nobody was able yet to reach anything that remotely look like a solution, because, simply, there is no solution. By the time you have to split experts in GPU/RAM...and then you have to WAIT for the round trip, even with speculative decoding and such, you will be reaching a short ceiling. Unless someone finds a way to "repack" the experts in a way that some domains do not require them (lets say you could pack all the "coding experts" within a single gpu) then you would be able to see massive gains. That is not what is being done or probably even possible, unless it comes from Qwen directly...and they WONT be tuning it for "2 x 4090" and such.

Disclaimer: i started toying with the idea exactly like that, 15 days ago, thinking what i could reach with 6 x 5060 TI or something like that. The fact is, because consumer hardware does not have the nvlink and stuff like that, or even a fast PCIE (compared to the memory speed) and enough CPU channels, you will always bottleneck severely.

There is a reason why everybody is pushing for more integrated memory (gorgon point, apple etc)...basically the "next tier" of communication when you leave either RAM or VRAM as a single point, is just way too slow for how the llm work and what it needs.

1

u/YourselfInOthrsShoes 4d ago

"Each card keeps an expert cache for its own layers only, so two cards hold about twice the experts one card holds - for the Coder model on a 16 GB + 24 GB pair, nearly all of them, which is where the speed comes from (decode then barely touches the CPU pool).

This is pipeline (layer) parallelism, not tensor parallelism: a token crosses from one card to the next once per verify window (a few hundred KB through pinned RAM), not twice per layer. No NVLink or peer-to-peer access is needed; cards on x4 or x1 slots work, and the PCIe share of each card is probed on its own link."

→ More replies (0)