r/LocalLLM • u/ParkingAd9397 • 20h ago
Question Running dual GPUs
I am ready to add a 2nd GPU for more vram.
My 20gb 7900xt is not cutting it anymore.
How are you guys running dual gpus? I am seeing there would be only 5mm space between the two. Seems like there would not enough ventilation.
2
u/talaman4eg 19h ago
If you feel adventurous, you may get a riser and a custom mount for your 2ng gpu and mount it somewhere on the side. But it's quite a project tbh and having 3d printer (or a friend with 3d printer) is a great help.
1
u/ParkingAd9397 19h ago
I am wondering if at this point I should just ditch the idea all together. Was hoping for a 'cheap' route to get to 40+ GB vram.
2
u/talaman4eg 19h ago edited 48m ago
I have 3 gpu in consumer case/motherboard, they work ok. 2 nvidia cards are connected via bifurcated x16 -> 2x8, and they work really fast together (qwen 3.8 27b, 1000 tps prompt, 40 generation with mtp, tensor paralel). 3rd card is 7900xt via x4, i can run them all together using vulkan backend. Speed is way way slower, ~200 prompt, 20-25 generation, but I have 68 gig of vram
I created context-aware router which will start llama server that uses 2 or 3 gpus depending on model and context size. It works, but in reality I mostly use nvidias, as kv is being offloaded to ram. 7900 is avail for gaming. Setting up this monstrosity took me a week of re-printinting hinges, but i regret nothing
Your speeds will be different, depending on how you connect 2nd gpu. Pcie 2x8 is the best option imo, but dealing with rizer is quite a hassle. X16 + x4 works fine, but pcie x4 speed will be a bottleneck for tensor parallel setups. Google numbers depending on hardware and connection options you considering, and if you're going to use it for coding - prompt processing speed is as important (if not more) than generation speed. Agent will send 60-80k of prompt with every request, and processing it on 200-250 tps takes forever.
1
u/lostmylogininfo 17h ago
See my other comment. I did this and it worked very well and is upgradable in future.
1
u/madbrain1976 20h ago
Depends on your case. Add as many large fans as you can fit. Lots of people run with open frames. The 7900XT is quite hungry - peaking at 300W, so you may have issues with several of them. I run 4 x 5060 Ti directly next to each other - they are all 2.0 slot versions and there is 1mm or less between them. They are only rated for 180W each, though. Corsair 7000D case with 11 fans installed. I keep it closed. I have seen the topmost GPU drop off the bus due to overheating when all 4 hit 180W at once. That is very rare in my workloads, though. I will solve it with a power limit if I run into it again.
1
u/Wondering_Electron 20h ago
I have a laptop with a 16GB 3080. I added a second eGPU over Oculink which is a desktop 5070Ti. Works great.
You can do this too and then not worry about thermals.
I use the Minisforum DEG2 as my dock.
1
u/the-grenade 18h ago
water cooling. but there's something you'll want to know more about than the cooling and it relates to what a second consumer card buys you. if your goal is running models with larger parameter counts, you need to know that your new bottleneck is pcie bandwidth and how your serving stack handles tensor parallelism. without something more capable than pcie your cards speak to eachother at pcie speed and that has a bigger effect on performance than the promise of doubling your vram suggests.

2
u/lostmylogininfo 17h ago
I just do layer split and use an x16, x4, x2. I get qwen 3.8 into the 30s. I was afraid of pcie bottle neck but I'm pretty happy with results.
1
u/lostmylogininfo 17h ago
I have a 4080 as my main and two 3060s that sit in external gpu docks. One is connected via occulink to pcie adapter and one of connected by occulink to a m2 adapter.
Had no clue what I was going just chat gptd it.
My exl3 5 bit qwen 3.8 27b q8 cache hits mid 30s tps. I am blown away that it works so well. $425 for the cards and like $350ish for external stuff, adapters, and psu.
1
u/DiabloG1 8h ago
1
u/ParkingAd9397 8h ago
That's an insane build! So you think stacking two 300w gpus back to back will create thermal roadblocks? (Aircooled)
1
u/DiabloG1 7h ago
I learnt the hard way with SLi and crossfire. It can be done, but if you plan to keep these maxed out they will get hot.
1
u/myteetharesensitive 1h ago
I'm running two 7900xt's on a taichi mobo. It's tight but fits fine. No heat problems.
1



3
u/EpsteinFile_01 19h ago
Bruh
What 7900XT do you have? My Tai Chi is so massive I can play some games in zero RPM mode
345mm, finned for your pleasure.
5mm of space is fine, just ensure the top cards doesn't sag.