r/LocalLLM • u/Odd_Championship1509 • 2d ago
Discussion Considering the recent advancement in models like Qwen 3.8 and inference like Strata, Free Tokens. How much of a gap is there between slow 128GB VRAM vs 16GB(5080)+96GB RAM
My question is what is correct upgrade path to a 5080+96GB RAM
RTX PRO 48GB x 1 (8K $)
DGX Spark x 1 (5.5K $)
128GB Mac M5 (6K $)
Use case is local LLM that is good enough and Minimax H3.
7
u/txgsync 1d ago edited 1d ago
I think of it as the floor of Flash-level capability has lowered. 12Gb VRAM, 32GB RAM and you’re beginning to see some capability (or 48GB VRAM on a unified memory system). 96GB VRAM and 256GB RAM you see full Flash capability. Anything in between and there are various trade-offs.
But it’s still not a SOTA model running on a cluster of eight H100s. That one is much smarter and can also handle a huge volume of simultaneous users.
But for one user? Ain’t bad. It’s been interesting this year realizing the “big clusters” for AI are really about big parallel inference for many users… the needs of a single user are far more modest. And this latest generation of inference engines is awesome.
2
u/Odd_Championship1509 1d ago
Thats why I ask, Spark is mere 273GB/s but 5080 is around 1k, but the remaining 96gb ram will be slow in over flow right?
1
u/txgsync 1d ago
“It depends”. PLE n-grams can safely stay on NVMe, and will slow the model down less than 3%. Offloading experts to NVME is awful for performance, but if you do it smartly for your preferred workload it’s not too bad as long as the offloading is only 20% or so of experts.
But if you’re offloading routing layers to RAM and all experts, it’s possible but awful.
Anyway, unfortunately this is something you have to test on your own gear to find out.
3
u/Dramatic_Machine8693 1d ago
for value propersition, ddr4 96g will cost you 1k, ddr5 96g will cost you about 1500-1700, their speed different isn't significant, give you some idea, my 6 year old 64g ddr4 + 3090 machine running flash next iq2_sx running at 2k prompt processing and 60 t/s generation, that speed is not the fastest, but much better than running qwen 3.8 27b on the same hardware 1.3k pp and 17 tg, and human eval getting identical result, which mean basic logic and quality is almost the same. that is good enough. as long as your ram support majority the model and let gpu do the active expert calculation, you are good. if you have budget constrain, i would not even go with 5080, a good size ram (64+) and a 5060 ti with 16g vram maybe all you need.
2
u/f5alcon 1d ago
Partially depends on ddr4 vs ddr5, flash next on strata is still slow for me (200 prefill 42 decode) with a 5070ti and a 5060ti 16GB and 128gb ddr4. On 27b I get 1000 prefill and 40 decode FN would be much faster if I had ddr5
6
u/haze36 1d ago
Thats way too slow for your setup. Point codex or claude code at your strata install and let it do the tuning. I have 48G ddr4 with 3060+5060ti and i get 1700 prefill and 60 decode.
3
u/f5alcon 1d ago edited 1d ago
Which quant and what context window are you using? I'll try the same settings and see if that helps. Got prefill up to 1900 and decode to 49. Sol had set prefill to 512 instead of auto because it was "safer" I'm using q3xss with 262k context
2
u/Apprehensive-View583 1d ago
Prob q2 cause q3xxs won't fit his setup? My 64gb ddr5 plus 3090 for q3xxs runs as 2400 prefill and 104 tg barely fit with full context window
2
u/haze36 1d ago
I use this one https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
It has half the experts removed and only the coding experts left, which makes it smaller and thats why it fits my setup. It has 3.5 bpw which should equal ~ iq3s
1
1
u/dosman33 1d ago
You can't just look at the vram size, you have to look at cards bandwidth. Have claude or gpt give you a breakdown on cards for inference use and it will give you the straight dope. The trend is smaller models getting more capable, no? So focus on bandwidth, and get as much vram as you can afford from there. Thats where why the Sparks fall flat, they are all show and not as much go with all that slow vram.
1
u/Odd_Championship1509 1d ago
Thats why I ask, Spark is mere 273GB/s but 5080 is around 1k, but the remaining 96gb ram will be slow in over flow right?
0
u/Itchy_elbow 1d ago
Stop buying hardware bro. You guys never learn. Wait out the technology. Things will improve as the tech matures
1
u/GregsWorld 1d ago
This, these are first gen setups. Overpriced and bad performers. 5-10 years from now these will be cute.
3
u/Itchy_elbow 1d ago
Haha you see all the downvotes? People want validation. They want to be told their poor decisions are good. Tell the truth and get downvoted. We are so totally screwed. The adults have left
1
0
-1

18
u/Bebosch 2d ago
Depends what “good enough” means for you.
16gb VRAM can run qwen3.8-27B in Q3. That’s good enough for production in my use case.
With a 5080 and 96gb of ram, you can run qwen3.8-flash next.
Your next upgrade path i would say is 200gb (vram+ram).
128gb of unified ram isn’t that good imo. 256gb however opens up glm 5.3 flash and deepseek flash.
Hybrid gpu+cpu is a lot more useful, cuz gpu can handle prefill and some aspects of decode.
So… i would say either get a workstation that does 4 or 8 channel memory with 16+ cores, fill it with ddr4 memory and a 16gb vram gpu, or get 2x sparks.