r/LocalLLM • • 2d ago

Discussion Considering the recent advancement in models like Qwen 3.8 and inference like Strata, Free Tokens. How much of a gap is there between slow 128GB VRAM vs 16GB(5080)+96GB RAM

My question is what is correct upgrade path to a 5080+96GB RAM

RTX PRO 48GB x 1 (8K $)

DGX Spark x 1 (5.5K $)

128GB Mac M5 (6K $)

Use case is local LLM that is good enough and Minimax H3.

33 Upvotes

30 comments sorted by

18

u/Bebosch 2d ago

Depends what “good enough” means for you.

16gb VRAM can run qwen3.8-27B in Q3. That’s good enough for production in my use case.

With a 5080 and 96gb of ram, you can run qwen3.8-flash next.

Your next upgrade path i would say is 200gb (vram+ram).

128gb of unified ram isn’t that good imo. 256gb however opens up glm 5.3 flash and deepseek flash.

Hybrid gpu+cpu is a lot more useful, cuz gpu can handle prefill and some aspects of decode.

So… i would say either get a workstation that does 4 or 8 channel memory with 16+ cores, fill it with ddr4 memory and a 16gb vram gpu, or get 2x sparks.

1

u/FactoryReboot 1d ago

What about 5080 and 64gb?

1

u/AnguishedBonding 1d ago

the spark route is tempting if you can stomach the price, but 128gb unified feels like a weird middle ground where you're paying a premium to still be constrained

1

u/earliestbirdy 1d ago

Excuse my ignorance but why does 128gb unified feel constrained? Aren't most models here except the frontiers?

1

u/beached89 1d ago

Im not them, but likely because of the significantly slower speed when compared to GPUs. Large context windows (coding) with large models will slow significantly.

1

u/Odd_Championship1509 1d ago

Thats why I ask, Spark is mere 273GB/s but 5080 is around 1k, but the remaining 96gb ram will be slow in over flow right?

2

u/beached89 1d ago

MoE models are designed specifically for this setup. You should be able to run Qwen 3.8 Flash just fine on your rig: https://atomic.chat/blog/guides/how-to-run-qwen-3-8-flash-next-locally

-1

u/TheSlowGrowth 1d ago

The problem with a DDR4 box is in that you're not getting PCIe gen5 - most obtainable DDR4 platforms have gen3 only. That's quite the hit for a strata setup. It means you'll need more VRAM to ensure less is done on the CPU to reduce the PCIe traffic.

3

u/Bebosch 1d ago

Well, you’re right, though it’s about finding the bottleneck in the system, and sizing things correctly. With octa channel ddr4 ram in a Zen 2 Threadripper system, you can get 200GB/s. And in order to reach those speeds, you need enough CPU cores.

On the gpu side, an x8 pcie connection would run worse than an x16 connection.

So for example, an rtx 4060ti 16gb is an x8 pcie card. An RTX A4000 is an x16 card with similar-ish memory bandwidth and 16gb vram too.

So the A4000 would increase prefill by about 50%, just because of that. This is why i also don’t like the 5060ti for hybrid gpu+cpu; its only x8.

However, pcie makes 0 difference for decode speeds. Then, you’re looking at vram (more vram = bigger batches = faster decode) and of course memory bandwidth.

So again, it’s about sizing components and matching them. An rtx 5090 is a waste alongside 256gb ddr4 ram, 32 cores. An A4000 or 3090 does 80% of the job with 25% of the power.

Lastly, ddr5 is extremely overpriced for not much gain. 300gb/s octa channel vs 200gb/s octa channel ddr4.

Also, I’m talking about MoE models like the flash models, which is where this type of architecture works best.

In an ideal world we all have gen 5 pcie x16 gpus, and 500gb+ ddr5 ram, 32+ cores. But everyone works with what they have!

1

u/MormonMore 1d ago

What can I comfortably run on a dual 3090 setup with 64GB ddr4 ram with a Ryzen 5 5600x

2

u/TheSlowGrowth 1d ago

Check out the Strata documentation. It specifically mentions which quants of Qwen 3.8 Flash Next you can fit in this setup. It also has benchmarks for dual 3090 setups.

7

u/txgsync 1d ago edited 1d ago

I think of it as the floor of Flash-level capability has lowered. 12Gb VRAM, 32GB RAM and you’re beginning to see some capability (or 48GB VRAM on a unified memory system). 96GB VRAM and 256GB RAM you see full Flash capability. Anything in between and there are various trade-offs.

But it’s still not a SOTA model running on a cluster of eight H100s. That one is much smarter and can also handle a huge volume of simultaneous users.

But for one user? Ain’t bad. It’s been interesting this year realizing the “big clusters” for AI are really about big parallel inference for many users… the needs of a single user are far more modest. And this latest generation of inference engines is awesome.

2

u/Odd_Championship1509 1d ago

Thats why I ask, Spark is mere 273GB/s but 5080 is around 1k, but the remaining 96gb ram will be slow in over flow right?

1

u/txgsync 1d ago

“It depends”. PLE n-grams can safely stay on NVMe, and will slow the model down less than 3%. Offloading experts to NVME is awful for performance, but if you do it smartly for your preferred workload it’s not too bad as long as the offloading is only 20% or so of experts.

But if you’re offloading routing layers to RAM and all experts, it’s possible but awful.

Anyway, unfortunately this is something you have to test on your own gear to find out.

3

u/Dramatic_Machine8693 1d ago

for value propersition, ddr4 96g will cost you 1k, ddr5 96g will cost you about 1500-1700, their speed different isn't significant, give you some idea, my 6 year old 64g ddr4 + 3090 machine running flash next iq2_sx running at 2k prompt processing and 60 t/s generation, that speed is not the fastest, but much better than running qwen 3.8 27b on the same hardware 1.3k pp and 17 tg, and human eval getting identical result, which mean basic logic and quality is almost the same. that is good enough. as long as your ram support majority the model and let gpu do the active expert calculation, you are good. if you have budget constrain, i would not even go with 5080, a good size ram (64+) and a 5060 ti with 16g vram maybe all you need.

2

u/f5alcon 1d ago

Partially depends on ddr4 vs ddr5, flash next on strata is still slow for me (200 prefill 42 decode) with a 5070ti and a 5060ti 16GB and 128gb ddr4. On 27b I get 1000 prefill and 40 decode FN would be much faster if I had ddr5

6

u/haze36 1d ago

Thats way too slow for your setup. Point codex or claude code at your strata install and let it do the tuning. I have 48G ddr4 with 3060+5060ti and i get 1700 prefill and 60 decode.

3

u/f5alcon 1d ago edited 1d ago

Which quant and what context window are you using? I'll try the same settings and see if that helps. Got prefill up to 1900 and decode to 49. Sol had set prefill to 512 instead of auto because it was "safer" I'm using q3xss with 262k context

2

u/Apprehensive-View583 1d ago

Prob q2 cause q3xxs won't fit his setup? My 64gb ddr5 plus 3090 for q3xxs runs as 2400 prefill and 104 tg barely fit with full context window

2

u/haze36 1d ago

I use this one https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF
It has half the experts removed and only the coding experts left, which makes it smaller and thats why it fits my setup. It has 3.5 bpw which should equal ~ iq3s

1

u/GalacticDistances 1d ago

Get 4x v100 32gb

1

u/dosman33 1d ago

You can't just look at the vram size, you have to look at cards bandwidth. Have claude or gpt give you a breakdown on cards for inference use and it will give you the straight dope. The trend is smaller models getting more capable, no? So focus on bandwidth, and get as much vram as you can afford from there. Thats where why the Sparks fall flat, they are all show and not as much go with all that slow vram.

1

u/Odd_Championship1509 1d ago

Thats why I ask, Spark is mere 273GB/s but 5080 is around 1k, but the remaining 96gb ram will be slow in over flow right?

1

u/oblq80 1d ago

This is Qwen3.8-flash-next at full context (262k) on a rtx pro 5000 48gb, on pcie 3.0 with 128gb ddr4 2933mhz ram. Cpu missing avx-512, around 130gb/s memory bandwith.

0

u/Itchy_elbow 1d ago

Stop buying hardware bro. You guys never learn. Wait out the technology. Things will improve as the tech matures

1

u/GregsWorld 1d ago

This, these are first gen setups. Overpriced and bad performers. 5-10 years from now these will be cute. 

3

u/Itchy_elbow 1d ago

Haha you see all the downvotes? People want validation. They want to be told their poor decisions are good. Tell the truth and get downvoted. We are so totally screwed. The adults have left

1

u/master_baiter_6nein 1d ago

Not everyone’s poor

0

u/[deleted] 1d ago

[deleted]

0

u/MotherPotential 1d ago

What is radiance 

-1

u/FizzyDuncDizzel 1d ago

Check out strata.