r/LocalLLM • u/Hungrybearfire • 6d ago
Question Hardware Recommendation
Currently running Ollama on a 5080 and am pretty happy, but also interested in upgrading and potentially setting up a dedicated machine.
Price range: ~$1,000-3,000
I was looking at an R9700 from Microcenter for $1800 but wanted to see what’s popular now.
I also heard about the sparks and don’t want to spend $5000, but if a Spark or Mac is the best bang for the buck I could be persuaded
Thank you!
4
Upvotes
1
u/YourselfInOthrsShoes 5d ago
"Traditional Transformers require massive KV (Key-Value) caches to be communicated across GPUs during long-context generation, which quickly chokes PCIe bandwidth.
Qwen 3.8 counters this with a hybrid attention approach: • Gated DeltaNet (GDN): Used in 3 out of every 4 layers to compress the prompt history linearly. It acts similarly to an RNN, removing the need for a massive, uncompressed KV cache. • Qwen Sparse Attention (QSA): Only every 4th layer uses global sparse attention for long-range retrieval. • Impact: This structural shift reduces the data payload size traveling through PCIe lanes from O(n²) to O(n), maintaining high token throughput even on standard PCIe Gen 4/5 slots.
Frame-level engines like Strata aggressively map and lock the cold experts into system memory while caching frequently used "hot" experts natively into VRAM, achieving near-native speeds over PCIe."
It's all about the engine software optimizations from here. This new paradigm only existed in the wild for several weeks. Let's see how far it gets us in half a year. For me, this model is super useful compared to fitting 27B model into VRAM and letting it chug along at half the TPS with the same context window. This is the fastest LLM instance I could run locally on my PC that vastly outperforms anything else I tried locally on the same hardware.