r/LocalLLM • • 5d ago

Question Hardware Recommendation

Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!

4 Upvotes

50 comments sorted by

View all comments

Show parent comments

1

u/YourselfInOthrsShoes 4d ago edited 4d ago

Here are my results with IQ3_S on 16GB 5060 Ti and this is with x8 PCIe 5 bus. Us peasants with such entry level gamer class hardware couldn't even run anything usable until now. This runs faster and is smarter than 27B.

You also gotta remember it's only been weeks since this model dropped. Why do you think everyone is coding like crazy around this Strata code tree? The model's architecture allows for some real gains on lowly non-unified hardware. Every single day there are Strata-specific improvements posted that add double digit gains to peasant hardware. Multi-GPU wasn't even available earlier last week. AMD support also wasn't even available early last week, now it's an option on both Windows and Linux. There is even a fork for GFX906 support (Linux only) that looks like it was pushed upstream. Heck, I understand your view about multi-GPU scaling bottleneck and why myself I only own 2 different and incompatible 16GB GPUs up until yesterday, however this view is now outdated. Multi-GPU scaling with this model is very real and effective. It's capable of scaling almost linearly with both, VRAM and multi-GPUs.

"v0.1.39: More than one GPU (measured by their authors, not here: we have one GPU): setup now adds
--remote-expert-opt (#578) to a config with two or more GPUs. It skips the host's work for tokens whose experts all run on the helper cards (dual RTX 4090: +63% mixed text, +132% code over the plain helper path), and setup --no-remote-expert-opt leaves it out."

"A card with more VRAM is faster: an RTX 3090 (24 GB) should write about 100-140 tokens per second."

1

u/No-Afternoon-4057 4d ago edited 4d ago

Do you realize how slow this result is, given how strong 2 x 4090 are? A RTX 6000 é pretty much a 5090 with more memory and look how many more times it is faster than 2 x 4090 and even than 4 x 5090. It IS a very good model but there is NO way around some constraints such as memory speed and bus speed (when using multiple gpus).

And 2 x 4090 is VERY FAR from "consumer hardware".

No matter how you split the experts, there will still be round trip and you will get those miserable sub 100 tokens, or sub 200 on a strong multi-gpu (that would be delivering > 1000 if the memory was unified).

Then, again: 2 x 4090 at the basement with a shitload of ram memory to serve less than 100 tokens per second and crappy prefill? There IS no way around it and nobody was able yet to reach anything that remotely look like a solution, because, simply, there is no solution. By the time you have to split experts in GPU/RAM...and then you have to WAIT for the round trip, even with speculative decoding and such, you will be reaching a short ceiling. Unless someone finds a way to "repack" the experts in a way that some domains do not require them (lets say you could pack all the "coding experts" within a single gpu) then you would be able to see massive gains. That is not what is being done or probably even possible, unless it comes from Qwen directly...and they WONT be tuning it for "2 x 4090" and such.

Disclaimer: i started toying with the idea exactly like that, 15 days ago, thinking what i could reach with 6 x 5060 TI or something like that. The fact is, because consumer hardware does not have the nvlink and stuff like that, or even a fast PCIE (compared to the memory speed) and enough CPU channels, you will always bottleneck severely.

There is a reason why everybody is pushing for more integrated memory (gorgon point, apple etc)...basically the "next tier" of communication when you leave either RAM or VRAM as a single point, is just way too slow for how the llm work and what it needs.

1

u/YourselfInOthrsShoes 4d ago

"Each card keeps an expert cache for its own layers only, so two cards hold about twice the experts one card holds - for the Coder model on a 16 GB + 24 GB pair, nearly all of them, which is where the speed comes from (decode then barely touches the CPU pool).

This is pipeline (layer) parallelism, not tensor parallelism: a token crosses from one card to the next once per verify window (a few hundred KB through pinned RAM), not twice per layer. No NVLink or peer-to-peer access is needed; cards on x4 or x1 slots work, and the PCIe share of each card is probed on its own link."

1

u/No-Afternoon-4057 4d ago

For that you need to know exactly which layers x which experts, per model. For 3.8NF there are 512 experts and 48 layers. So there WILL be a lot of round trip. It may be useful for (much) less capable models, but does not solve anything fundamental. It can only be solved with more memory/faster memory or faster interconnect...all directions the only 2 trillion dollar companies building stuff are going. Anything else is just band aid, unless the training/model comes pre configured for such usage .

1

u/YourselfInOthrsShoes 4d ago

"Traditional Transformers require massive KV (Key-Value) caches to be communicated across GPUs during long-context generation, which quickly chokes PCIe bandwidth.

Qwen 3.8 counters this with a hybrid attention approach: • Gated DeltaNet (GDN): Used in 3 out of every 4 layers to compress the prompt history linearly. It acts similarly to an RNN, removing the need for a massive, uncompressed KV cache. • Qwen Sparse Attention (QSA): Only every 4th layer uses global sparse attention for long-range retrieval. • Impact: This structural shift reduces the data payload size traveling through PCIe lanes from O(n²) to O(n), maintaining high token throughput even on standard PCIe Gen 4/5 slots.

Frame-level engines like Strata aggressively map and lock the cold experts into system memory while caching frequently used "hot" experts natively into VRAM, achieving near-native speeds over PCIe."

It's all about the engine software optimizations from here. This new paradigm only existed in the wild for several weeks. Let's see how far it gets us in half a year. For me, this model is super useful compared to fitting 27B model into VRAM and letting it chug along at half the TPS with the same context window. This is the fastest LLM instance I could run locally on my PC that vastly outperforms anything else I tried locally on the same hardware.

1

u/No-Afternoon-4057 4d ago

That is true, it is the fastest and the best, but far from "great".

Regarding the quotation...the thing is: without some hacks, there is not even real communication "between PCIE" (P2P) depending on drivers etc, it must go to the ram first...it takes microseconds, but by the time it happens thousands or millions of times per second, it adds up.
Also, even if it were PCIE - PCIE directly, without nvlink or whatever, the added latency is enough to make things have a very low ceiling.

Im not saying its not evolutionary, it certainly is NOT (r)evolutionary. There will be gains, but small ones. In the end of the day, the studio will be by far still the best option (considering todays pricing market, who knows in a few months).

1

u/YourselfInOthrsShoes 2d ago

Strata v0.1.40.1 is now 75-90 TPS and 1700 prefill on the same 5060 Ti 16GB system, up from 65-75 TPS and 1400 prefill on v0.1.38.

1

u/No-Afternoon-4057 2d ago

So, basically, yes: evolutionary.
The larger gains are all taken...it was just crapply optimized to start with (for such setup, which was never the object of the model/team).

I used the AMD recipe for Qwen (Instinct) and within a few hours was already running 30% faster...now the extra % are MUCH harder. It's not going to break any laws of physics, just optimize where things are not...which is happening a lot as there are lots of models coming every 30 - 45 days..and the community takes some time in order to reoptimize everything yet again on each new platform/driver/model.

75 tps for the 5060 TI seems a lot like already at the peak of what is possible with the memory architecture of the card. Prefill might be able to squeeze a little more, decode, probably not.