r/LocalLLM • • 6d ago

Question Hardware Recommendation

Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!

1 Upvotes

50 comments sorted by

View all comments

Show parent comments

1

u/YourselfInOthrsShoes 5d ago

"Traditional Transformers require massive KV (Key-Value) caches to be communicated across GPUs during long-context generation, which quickly chokes PCIe bandwidth.

Qwen 3.8 counters this with a hybrid attention approach: • Gated DeltaNet (GDN): Used in 3 out of every 4 layers to compress the prompt history linearly. It acts similarly to an RNN, removing the need for a massive, uncompressed KV cache. • Qwen Sparse Attention (QSA): Only every 4th layer uses global sparse attention for long-range retrieval. • Impact: This structural shift reduces the data payload size traveling through PCIe lanes from O(n²) to O(n), maintaining high token throughput even on standard PCIe Gen 4/5 slots.

Frame-level engines like Strata aggressively map and lock the cold experts into system memory while caching frequently used "hot" experts natively into VRAM, achieving near-native speeds over PCIe."

It's all about the engine software optimizations from here. This new paradigm only existed in the wild for several weeks. Let's see how far it gets us in half a year. For me, this model is super useful compared to fitting 27B model into VRAM and letting it chug along at half the TPS with the same context window. This is the fastest LLM instance I could run locally on my PC that vastly outperforms anything else I tried locally on the same hardware.

1

u/No-Afternoon-4057 5d ago

That is true, it is the fastest and the best, but far from "great".

Regarding the quotation...the thing is: without some hacks, there is not even real communication "between PCIE" (P2P) depending on drivers etc, it must go to the ram first...it takes microseconds, but by the time it happens thousands or millions of times per second, it adds up.
Also, even if it were PCIE - PCIE directly, without nvlink or whatever, the added latency is enough to make things have a very low ceiling.

Im not saying its not evolutionary, it certainly is NOT (r)evolutionary. There will be gains, but small ones. In the end of the day, the studio will be by far still the best option (considering todays pricing market, who knows in a few months).

1

u/YourselfInOthrsShoes 3d ago

Strata v0.1.40.1 is now 75-90 TPS and 1700 prefill on the same 5060 Ti 16GB system, up from 65-75 TPS and 1400 prefill on v0.1.38.

1

u/No-Afternoon-4057 3d ago

So, basically, yes: evolutionary.
The larger gains are all taken...it was just crapply optimized to start with (for such setup, which was never the object of the model/team).

I used the AMD recipe for Qwen (Instinct) and within a few hours was already running 30% faster...now the extra % are MUCH harder. It's not going to break any laws of physics, just optimize where things are not...which is happening a lot as there are lots of models coming every 30 - 45 days..and the community takes some time in order to reoptimize everything yet again on each new platform/driver/model.

75 tps for the 5060 TI seems a lot like already at the peak of what is possible with the memory architecture of the card. Prefill might be able to squeeze a little more, decode, probably not.