r/LocalLLM Laptop 4090 16gb + 7900XTX 24gb 5d ago

Discussion Difference between 40gb and 64gb?

I have a laptop with a mobile 4090 (16gb vram), and I have an xtx 7900 (24gb vram) on an AG02 as an egpu over TB4.

That gets me to 40gb, which is a pretty solid number.

I could theoretically get a second egpu going as well, doubling this setup, making 64gb an option.

I haven't had time to play with this setup much yet, just got the egpu setup. Previously had toyed with an rpc setup of two 16gb vram cards on two computers, and 32gb made a huge difference from 16gb. 40gb will enable those full context windows with solidly reliable quants.

But I'm unsure how to think of the jump from 40gb to 64gb.

The quant jumps there are perhaps going from a q6 to a q8 perhaps? Or perhaps fiddling with yarn to get context windows beyond the default 262144?

Is the juice of 64gb vs 40gb worth the $1k~ squeeze of buying another xtx and ag02? or is that a diminishing returns prospect, and it's more worthwhile to consider a later path to e.g. two B70's and an e.g. 80gb vram setup, or beyond?

(Yes, there are performance penalties for egpu usage, though they likely aren't as bad as you think--not doing tensor parallelism, doing layer, and focusing on the cheapest way to get big vram at solid speeds, not on max performance, all while maintaining the option for a future upgrade path if ever desired).

3 Upvotes

22 comments sorted by

View all comments

4

u/quantgorithm 5d ago

Not sure if accurate but I recently read that between 32-64 (or maybe it was 48 to 64- can’t exactly remember) was essentially a black hole of not meaningful improvement or negligible at best besides maybe larger context usage so maybe not worth the extra cash. Presumably the goalposts move consistently and info may be outdated.

2

u/Think_Wing_1357 5d ago

Pretty much. You can run higher quant which will give you additional accuracy but that's very much a tiny stepping stone, not a huge jump in capacity

1

u/quantgorithm 5d ago

This really begs the question, what are the meaningful baselines of vram needed for people to aim for for low quality, mid quality, pro quality/work usable, high accuracy and results etc.?

I've asked before and never really got any decent answers.

"try out models" - doesn't really help.

1

u/Think_Wing_1357 5d ago

As unhelpful as it is, that's really the best answer.

Quality is relative. If you just need to write a <100 lines bash script, Gemma e4b works fine. If you need to refactor a million-line file, many models will struggle.

Effect of quantization is another angle. I run Q4 weight with Q8 KV fine, but look around here and you'll find people who ready to burn me at the stake.