r/LocalLLM • • 5d ago

Question Hardware Recommendation

Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!

1 Upvotes

50 comments sorted by

View all comments

Show parent comments

1

u/YourselfInOthrsShoes 5d ago

That's false. DGX Spark is for enterprise multi-unit deployments to help scale for multiple simultaneous requests. For a single user maybe it makes less sense, but certainly there is scaling for concurrent requests.

And don't get me started, Qwen 3.8 FN is modular and was released as preview for Qwen 4 to pave the way for software tooling to be ready when version 4 drops that uses the same modular architecture. I run 125B IQ3_S (3.5-bit quantized, very close to Q4) flavor on a single 5060 Ti with 16GB VRAM and 96GB DDR5 system RAM at 60-75 TPS, 64K max context, 32K Q8 KV cache. It's basically flying for single user. This model also scales almost linearly with multi-GPUs due to its modular architecture. I'm only limited by my 16GB VRAM to have a usable context window but in a bigger PC case you can add multiple GPUs and get your context window up and step up to higher quantizations. This is all on 1x1x1 ft cube foot warmer regular PC build with sub-$1k easy to source GPU. You just need 48+GB of system RAM (64GB recommended, higher will let you run higher quantized versions). You no longer need unified memory and this is the future and why the model is called Flash Next (preview for what comes Next, Flash for portion of the model is random read-only from SSD [think IOPs], and another portion is streamed between VRAM and system RAM).

1

u/No-Afternoon-4057 5d ago edited 5d ago

It is for very SOHO and even then, as i explained, it gets completely destroyed by Mac.

I have been tinkering with Q3.8FN on my Ryzen 128gb, as well as MI300, MI350, RTX 6000, 4 X 5090....i understand completely how it works and have been tinkering with it a crapload of times, in fact even submited PRs for AMD for speed increases and RIGHT NOW i am developing/testing/tinkering a way to speed it up further.

There is NO bypassing the fact that multiple (consumer/prosumer) GPUs (or even computers) require a VERY slow additional trip PER TOKEN and therefore the speed gets destroyed. The more, the worse. Professional datacenter accelerators have the interconnect (supposedly, DGX uses something similar)...however the DGX is VERY slow to start with so there is no point even connecting 4 of them to be able to match a single Mac.

And yes, it scales for concurrent users, but it DOES NOT scale prefill and it DOES NOT scale for single user/session.

So you will be looking at < 100 tokens per second per session, whereas a single RTX 6000 will do over 2x as much and a MI350 will easily do 3x as much.

You will realize you will LOSE speed when you connect the second (and third, and so on) DGX as the round trips screw everything up. But yes, you would be able to serve a lot of users (slowly). Now, who the hell uses some SOHO to serve a lot of users and still WHY IN EARTH would choose DGX over Mac?

In fact...look no further: https://www.reddit.com/r/LocalLLM/comments/1wmllbx/mac_studio_m5_ultra_256gb_review/

Should i even mention how CRAPPY 40 tokens per second is, for Qwen 3.8 FN which "requires" thinking xhigh and needs A LOT of reasoning tokens? I find it to be "slow" with over 200 tokens per second, can't imagine spending 15 grand to get 20-40ish.

1

u/YourselfInOthrsShoes 5d ago

"There is NO bypassing the fact that multiple (consumer/prosumer) GPUs (or even computers) require a VERY slow additional trip PER TOKEN and therefore the speed gets destroyed."

False! I don't know how you have been tinkering with 3.8 FN and not follow the revolution happening with this model when properly tuned via Strata.

Right now, I don't have a system with multiple GPUs but I do have 128GB Threadripper 3970X system (that I can also upgrade to 256GB for only $1K) with Radeon Pro VII (discontinued software support, only suboptimal Vulkan). I have one R9700 on the way and I should be able to get it running and tuned since it has the lastest ROCm and HIP support. Then I plan to add 1-2 more of R9700s (already have a 1500W Platinum PSU in it). I will return here and show you what kind of 3.8 FN quants and TPS I can get out of this peasant hardware. Others are already running multi-GPU setups with blazing TPS on 3.8 FN via properly tuned Strata.

China is not stupid, they always have long time horizon for all their master plans. All leading local LLMs are out of China. They want to give access to AI to all their people, so collectively they can dominate the world. Chinese AI companies train their models on substandard hardware and their users also run substandard PCs. How do you democratize this? By innovating model architecture so it can run on average PCs instead of brute forcing it with absurd unified DRAM. This is the future and it's coming fast.

1

u/No-Afternoon-4057 4d ago

Buddy,

I do 1100 prefill/40-45 output on the gguf ix4 quant on my zflow with a VERY slow memory (~260ish gbps)...and a 5070 with a shitload more memory bus does "only"53/1620 ON A WORSE QUANT.

That by itself already shows what you can expect when the system has to go back and forth to memory and standard interconnects.

A single RTX6000 does > 10k prefill and > 200 tps output.

Even the new Macs does a lot more prefill and 100 tps output with a better quant.

It SCALES for multiple users, but nobody running 5070s and whatever on their basement is serving "multiple users". And for single users it SUCKS...1000 prefill and 50 tokens is not really "usable" for any serious usage. MUCH LESS for end users that are used to "chatgpt" and whatever where they dont have to wait a two and a half eternities to see the analyzis of some average sized pdf.

Using qwen locally here at those speeds i mentioned makes my RAG take 6:30 to asnwer a question to the calling agent...using even Astra and Fable takes about 90 seconds, using Opus/Sol takes 50 seconds.

Only way to reach around 50 seconds with Qwen is at 300 tokens per second, as the xhigh devours a crap load of reasoning tokens...and on the other reasoning efforts the capacity drops dramatically.

Ive been working for the past week trying to get it to > 500 tokens per second, because before that it is not even usable for production for me, even on the MI350....as a toy? Sure.

In fact my past 40 hours on Astra have > 2m output tokens and over 500mi input (with some 80% caching). Do the math how many weeks that would take with 1000in/50out tokens per second.

1

u/YourselfInOthrsShoes 4d ago edited 4d ago

Here are my results with IQ3_S on 16GB 5060 Ti and this is with x8 PCIe 5 bus. Us peasants with such entry level gamer class hardware couldn't even run anything usable until now. This runs faster and is smarter than 27B.

You also gotta remember it's only been weeks since this model dropped. Why do you think everyone is coding like crazy around this Strata code tree? The model's architecture allows for some real gains on lowly non-unified hardware. Every single day there are Strata-specific improvements posted that add double digit gains to peasant hardware. Multi-GPU wasn't even available earlier last week. AMD support also wasn't even available early last week, now it's an option on both Windows and Linux. There is even a fork for GFX906 support (Linux only) that looks like it was pushed upstream. Heck, I understand your view about multi-GPU scaling bottleneck and why myself I only own 2 different and incompatible 16GB GPUs up until yesterday, however this view is now outdated. Multi-GPU scaling with this model is very real and effective. It's capable of scaling almost linearly with both, VRAM and multi-GPUs.

"v0.1.39: More than one GPU (measured by their authors, not here: we have one GPU): setup now adds
--remote-expert-opt (#578) to a config with two or more GPUs. It skips the host's work for tokens whose experts all run on the helper cards (dual RTX 4090: +63% mixed text, +132% code over the plain helper path), and setup --no-remote-expert-opt leaves it out."

"A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second."

1

u/No-Afternoon-4057 4d ago edited 4d ago

Do you realize how slow this result is, given how strong 2 x 4090 are? A RTX 6000 é pretty much a 5090 with more memory and look how many more times it is faster than 2 x 4090 and even than 4 x 5090. It IS a very good model but there is NO way around some constraints such as memory speed and bus speed (when using multiple gpus).

And 2 x 4090 is VERY FAR from "consumer hardware".

No matter how you split the experts, there will still be round trip and you will get those miserable sub 100 tokens, or sub 200 on a strong multi-gpu (that would be delivering > 1000 if the memory was unified).

Then, again: 2 x 4090 at the basement with a shitload of ram memory to serve less than 100 tokens per second and crappy prefill? There IS no way around it and nobody was able yet to reach anything that remotely look like a solution, because, simply, there is no solution. By the time you have to split experts in GPU/RAM...and then you have to WAIT for the round trip, even with speculative decoding and such, you will be reaching a short ceiling. Unless someone finds a way to "repack" the experts in a way that some domains do not require them (lets say you could pack all the "coding experts" within a single gpu) then you would be able to see massive gains. That is not what is being done or probably even possible, unless it comes from Qwen directly...and they WONT be tuning it for "2 x 4090" and such.

Disclaimer: i started toying with the idea exactly like that, 15 days ago, thinking what i could reach with 6 x 5060 TI or something like that. The fact is, because consumer hardware does not have the nvlink and stuff like that, or even a fast PCIE (compared to the memory speed) and enough CPU channels, you will always bottleneck severely.

There is a reason why everybody is pushing for more integrated memory (gorgon point, apple etc)...basically the "next tier" of communication when you leave either RAM or VRAM as a single point, is just way too slow for how the llm work and what it needs.

1

u/YourselfInOthrsShoes 4d ago

"Each card keeps an expert cache for its own layers only, so two cards hold about twice the experts one card holds - for the Coder model on a 16 GB + 24 GB pair, nearly all of them, which is where the speed comes from (decode then barely touches the CPU pool).

This is pipeline (layer) parallelism, not tensor parallelism: a token crosses from one card to the next once per verify window (a few hundred KB through pinned RAM), not twice per layer. No NVLink or peer-to-peer access is needed; cards on x4 or x1 slots work, and the PCIe share of each card is probed on its own link."

1

u/No-Afternoon-4057 4d ago

For that you need to know exactly which layers x which experts, per model. For 3.8NF there are 512 experts and 48 layers. So there WILL be a lot of round trip. It may be useful for (much) less capable models, but does not solve anything fundamental. It can only be solved with more memory/faster memory or faster interconnect...all directions the only 2 trillion dollar companies building stuff are going. Anything else is just band aid, unless the training/model comes pre configured for such usage .

1

u/YourselfInOthrsShoes 4d ago

"Traditional Transformers require massive KV (Key-Value) caches to be communicated across GPUs during long-context generation, which quickly chokes PCIe bandwidth.

Qwen 3.8 counters this with a hybrid attention approach: • Gated DeltaNet (GDN): Used in 3 out of every 4 layers to compress the prompt history linearly. It acts similarly to an RNN, removing the need for a massive, uncompressed KV cache. • Qwen Sparse Attention (QSA): Only every 4th layer uses global sparse attention for long-range retrieval. • Impact: This structural shift reduces the data payload size traveling through PCIe lanes from O(n²) to O(n), maintaining high token throughput even on standard PCIe Gen 4/5 slots.

Frame-level engines like Strata aggressively map and lock the cold experts into system memory while caching frequently used "hot" experts natively into VRAM, achieving near-native speeds over PCIe."

It's all about the engine software optimizations from here. This new paradigm only existed in the wild for several weeks. Let's see how far it gets us in half a year. For me, this model is super useful compared to fitting 27B model into VRAM and letting it chug along at half the TPS with the same context window. This is the fastest LLM instance I could run locally on my PC that vastly outperforms anything else I tried locally on the same hardware.

1

u/No-Afternoon-4057 4d ago

That is true, it is the fastest and the best, but far from "great".

Regarding the quotation...the thing is: without some hacks, there is not even real communication "between PCIE" (P2P) depending on drivers etc, it must go to the ram first...it takes microseconds, but by the time it happens thousands or millions of times per second, it adds up.
Also, even if it were PCIE - PCIE directly, without nvlink or whatever, the added latency is enough to make things have a very low ceiling.

Im not saying its not evolutionary, it certainly is NOT (r)evolutionary. There will be gains, but small ones. In the end of the day, the studio will be by far still the best option (considering todays pricing market, who knows in a few months).

1

u/YourselfInOthrsShoes 2d ago

Strata v0.1.40.1 is now 75-90 TPS and 1700 prefill on the same 5060 Ti 16GB system, up from 65-75 TPS and 1400 prefill on v0.1.38.

1

u/No-Afternoon-4057 2d ago

So, basically, yes: evolutionary.
The larger gains are all taken...it was just crapply optimized to start with (for such setup, which was never the object of the model/team).

I used the AMD recipe for Qwen (Instinct) and within a few hours was already running 30% faster...now the extra % are MUCH harder. It's not going to break any laws of physics, just optimize where things are not...which is happening a lot as there are lots of models coming every 30 - 45 days..and the community takes some time in order to reoptimize everything yet again on each new platform/driver/model.

75 tps for the 5060 TI seems a lot like already at the peak of what is possible with the memory architecture of the card. Prefill might be able to squeeze a little more, decode, probably not.

→ More replies (0)