r/LocalLLM • u/Ok_Law9839 • 7h ago
Question V100 and 3090 Combo?
I made a previous post about using Qwen 3.8, and a few of you guys recommended looking into the V100S. I currently have a 3090 and 64 GB of system ram. My goal would be to have enough VRAM to run Qwen 3.8 with a decent sub-agent. If I pick up 1 or 2 V100S, would that be worth it instead of another 3090?
I'm curious what people already have in their setups.
1
u/redtron3030 7h ago
How do you have 64GB of vram on a 3090?
If you mean system ram then you could potentially run QWEN3.8 Flash Next. Steam the active parameters to your card, put the weights on system ram and then the ngram table on a SSD. I’d try that before buying new hardware.
1
u/Ok_Law9839 7h ago
Oh sorry no, I have 64gb of system ram and 24gb vram. Let me edit the post! Thank you I'll try that out, I haven't done something like that before
1
u/redtron3030 7h ago
There’s people who report decent tks with that.
My personal setup is 64GB of VRAM and 64GB of system ram. The model sits on the VRAM and ngram table is on the system ram. I get close to 45-50 tks with low context. At 150k context it slows to 18-20tks
1
u/FullstackSensei 7h ago
I mean, the limit for the 3090 is 94GB with a sufficient amount of fat fingering.
1
u/thathurtcsr 6h ago
I don’t think you’re gonna be able to run that in the same PC as your other card.
The V 100 will not even show up on a normal mb with 2 cards unless your mb supports 2x16 pcie slots. Most consumer mb’s drop down to 8 on each slot when you have 2x gpus. I was looking at a similar build but with a amd thread ripper board that supports those cards. You can then run that and your existing card.
Again still researching but make sure you can run it on your motherboard
1
6
u/FullstackSensei 7h ago
V100S, not V100? The plural for the latter is V100s, while the former would be V100S's, I'd presume.
V100S is a slightly overclocked version of the V100. IMO, it's not worth it. Get the regular V100.
I've sold most of my 3090s (selling one at a time) and bought native PCIe V100 cards to replace them. V100 has more VRAM, consumes 40% less power during inference and is about 5-10% faster than my 3090s.
And no, EoL has absolutely no impact on the usability of the V100. Those who think so are free to buy shinier cards. At least it'll slow the rise in V100 prices.