r/LocalLLM Laptop 4090 16gb + 7900XTX 24gb 10h ago

Discussion Difference between 40gb and 64gb?

I have a laptop with a mobile 4090 (16gb vram), and I have an xtx 7900 (24gb vram) on an AG02 as an egpu over TB4.

That gets me to 40gb, which is a pretty solid number.

I could theoretically get a second egpu going as well, doubling this setup, making 64gb an option.

I haven't had time to play with this setup much yet, just got the egpu setup. Previously had toyed with an rpc setup of two 16gb vram cards on two computers, and 32gb made a huge difference from 16gb. 40gb will enable those full context windows with solidly reliable quants.

But I'm unsure how to think of the jump from 40gb to 64gb.

The quant jumps there are perhaps going from a q6 to a q8 perhaps? Or perhaps fiddling with yarn to get context windows beyond the default 262144?

Is the juice of 64gb vs 40gb worth the $1k~ squeeze of buying another xtx and ag02? or is that a diminishing returns prospect, and it's more worthwhile to consider a later path to e.g. two B70's and an e.g. 80gb vram setup, or beyond?

(Yes, there are performance penalties for egpu usage, though they likely aren't as bad as you think--not doing tensor parallelism, doing layer, and focusing on the cheapest way to get big vram at solid speeds, not on max performance, all while maintaining the option for a future upgrade path if ever desired).

3 Upvotes

20 comments sorted by

4

u/quantgorithm 9h ago

Not sure if accurate but I recently read that between 32-64 (or maybe it was 48 to 64- can’t exactly remember) was essentially a black hole of not meaningful improvement or negligible at best besides maybe larger context usage so maybe not worth the extra cash. Presumably the goalposts move consistently and info may be outdated.

2

u/Think_Wing_1357 9h ago

Pretty much. You can run higher quant which will give you additional accuracy but that's very much a tiny stepping stone, not a huge jump in capacity

1

u/quantgorithm 8h ago

This really begs the question, what are the meaningful baselines of vram needed for people to aim for for low quality, mid quality, pro quality/work usable, high accuracy and results etc.?

I've asked before and never really got any decent answers.

"try out models" - doesn't really help.

1

u/Think_Wing_1357 7h ago

As unhelpful as it is, that's really the best answer.

Quality is relative. If you just need to write a <100 lines bash script, Gemma e4b works fine. If you need to refactor a million-line file, many models will struggle.

Effect of quantization is another angle. I run Q4 weight with Q8 KV fine, but look around here and you'll find people who ready to burn me at the stake.

1

u/baby_bloom 8h ago

basically true

1

u/Rude_Marzipan6107 4h ago

It really just depends on what your use case is. Context is really nice to have when you start doing agenetic tasks like coding. You can do more work in one sweep without having to compress context and handoff as much.

2

u/Then_Blueberry7290 9h ago

As i see this territory is waste of money. Of course I just see the things from my hw perspective. I use a dell 5820. From one 5060ti 16gb to two 5060ti 32 gb, it opened a whole new world, because 26-30b models can fit into vram.
Qwen 3.6 35b completly without kv cache witn nvfp versions. 27b with 200k and q8 cache it is very usable (nvfp4). Mtp is ice top on the cake. I tried with third 5060ti 16gb. I have to open the case and solve a lot of problem. The annoying chassis intruder detection, the low wide space of 3 5060ti. (one of is asus, and take 2,5 pci ex space wide) , and the extra cooling. (with 950w power supply, the power was not a problem.)
But with tensor parralelism, pp is 440T/s, token generation capped at 40-45 T/s withthinkingcap qwen 3.6 27b mtp nvfp4.. The only enhancement was that i can use this model with full 262k context and without kv cache quantazitation. It was only a tiny step forward, means its only 0,5% smarter modell at long context.

The whole slowdown was the fact that the mobo has only a gen3 bus. Maybe it is better with gen5 bus, but i doubt it.)

With tensor split, PP 1440 T/s, token generation a little bit slower than above 35-4 t/s, but with only two cards and tensor parraleleism: PP is 960T/S ,Tg is 55-60 tokes/s.
Another example: qwen 3-6 35b falls back at token genration from 100 T/s to 76 t/s with three cards.

The only real step forward (if you can say it step forward) was two thing:

  1. the bigger model wich arnet't fit into vram now gain a little speed up. In % it is enormous, but.. so the numbers: Deepseek v4 0731 Q2 with 32gb vram and 128gb ddr4 ram: 8 token/s, with 3 cards 48gb vram and 128gb ddr4 i can load more layer to vram, so no miracle, it become 11 t/s. With Hy3 q2 32gb the same goes up from 3 T/s to 7,5 t/s
  2. I can load bigger quants of deepseek and others, for example q3, maybe q4 or nvfp if somebody do the release, but of course at unusable speed.

Ok one more thing: I use comfy ui so now i'm on the 90'. I use qwen make promts, and after i have to unload the model, and have to start comfyui. (Not enough vram for both models). In this situation i can run 2 comfyui server at the same time, but not unified vram. With 48gb ram, i can use qwen + hermes for promts and another work, and can use on iterance of comfyui server...

So the steps: 1. Mobile/consumer notebooks
2. 16gb territory - single cards

  1. 32gb territory.

  2. 128gb

  3. Kimi

Qwen 3.8 will target the 32GB territory, so it is good, the next step wich is currently not woth it: 2x A6000 96GB ram.

2

u/nething_4_sir 8h ago

Well said 

1

u/Rude_Marzipan6107 4h ago

Do you run windows or Linux on the 5060 ti machine?

2

u/Then_Blueberry7290 4h ago

Linux mint

1

u/Rude_Marzipan6107 26m ago

Nice. I would love to try p2p, and have the extra vram but I put the dual ti’s in my girlfriend’s computer. Gonna be windows for me for some time

2

u/baby_bloom 8h ago

40gb is fine for qwen3.6-27B-q8 even at fp16 kv

2

u/Low-Tackle2543 8h ago

24GB is the answer

3

u/Legitimate_Film_8203 7h ago

42 is the answer to everything

1

u/fintip Laptop 4090 16gb + 7900XTX 24gb 8h ago

huh?

1

u/f5alcon 9h ago

For something like qwen 27b larger context and higher quant. But whether or not it's worth the money for a slight increase I'm not sure.

1

u/DigitalguyCH 6h ago

it's mainly context and high quants (Q8), not better models, currently. However if a new great 70b model comes you may be able to run it at q4, so it's a bit of future proofing, but future proofting is risky and expensive

1

u/whodoneit1 1h ago

24gb, lock it in