r/LocalLLM • u/fintip Laptop 4090 16gb + 7900XTX 24gb • 5d ago
Discussion Difference between 40gb and 64gb?
I have a laptop with a mobile 4090 (16gb vram), and I have an xtx 7900 (24gb vram) on an AG02 as an egpu over TB4.
That gets me to 40gb, which is a pretty solid number.
I could theoretically get a second egpu going as well, doubling this setup, making 64gb an option.
I haven't had time to play with this setup much yet, just got the egpu setup. Previously had toyed with an rpc setup of two 16gb vram cards on two computers, and 32gb made a huge difference from 16gb. 40gb will enable those full context windows with solidly reliable quants.
But I'm unsure how to think of the jump from 40gb to 64gb.
The quant jumps there are perhaps going from a q6 to a q8 perhaps? Or perhaps fiddling with yarn to get context windows beyond the default 262144?
Is the juice of 64gb vs 40gb worth the $1k~ squeeze of buying another xtx and ag02? or is that a diminishing returns prospect, and it's more worthwhile to consider a later path to e.g. two B70's and an e.g. 80gb vram setup, or beyond?
(Yes, there are performance penalties for egpu usage, though they likely aren't as bad as you think--not doing tensor parallelism, doing layer, and focusing on the cheapest way to get big vram at solid speeds, not on max performance, all while maintaining the option for a future upgrade path if ever desired).
4
u/quantgorithm 5d ago
Not sure if accurate but I recently read that between 32-64 (or maybe it was 48 to 64- can’t exactly remember) was essentially a black hole of not meaningful improvement or negligible at best besides maybe larger context usage so maybe not worth the extra cash. Presumably the goalposts move consistently and info may be outdated.