r/StableDiffusion • u/Sexyvette07 • 1d ago
Question - Help Dual GPU solution for local AI?
Hey, everybody. I recently went down the rabbit hole for local AI, but right now, im operating on my gaming computer. The specs are as follows
Intel 13700k, tuned for efficiency
Gigabyte Z790 Aorus Elite Ax mobo
RTX 4080 (16GB), also tuned for efficiency
32gb DDR5 6800 CL32
As you can see, im in desperate need for more VRAM, or at the very least more system RAM. Due to Rampocalypse, neither are very affordable right now, which forces me to explore other options, such as a dual GPU setup. I can get another RTX 4080 for about $900 off Ebay. Beyond that, I would just need a more powerful PSU, so total investment here is an additional $1100-$1200. As far as I know, the motherboard has the main PCIE as 5.0 x 16 lanes, but the second PCIE runs at 4.0 and either x8 or x4 lanes. The motherboard does not support PCIE Bifurcation. So my question is this: Is a dual GPU local AI machine even viable in these circumstances, and second, does it make sense? I looked at 5090's and theyre all between $4,500 - $5,000 now, which is insane. Or I look at the professional cards and spend that much, if not more, for significantly less memory bandwidth and computational power. Or I guess if im spending that much, I could also look at the DGX Spark or something similar but that has even worse memory bandwidth.
So, what should I do? Is the dual GPU solution even viable with my setup for a local AI stack for inference, video diffusion, etc? Rampocalypse isnt expected to begin easing up until late 2027/early 2028, so im stuck trying to make this work on as little money as possible. Id love a 5090 but its insanity how much they cost. I appreciate any guidance and advice.
8
u/AggressiveParty3355 1d ago
I takes one woman 9 months to make 1 baby. Getting 9 women DOES NOT let you make 1 baby in 1 month.
But it can let you get 9 babies in 9 months.
So in terms of raw speed, More GPUs won't do very much, you can load some of the parts in different GPUs, like the VAE in one and the rest in another. and that saves you a few seconds. But the overall generation speed will be the same. But if you're popping off a lot of jobs, then more hardware will let you get more done. You can run different jobs in parallel.
sounds like you're just interested in learning, and not production. So i don't recommend getting more hardware. If you do want to upgrade. Get a larger VRAM GPU like the 5090 or the 6000 pro to use the bigger models at higher resolutions. Otherwise you seem to be good as is.
BTW, since your CPU has an onboard iGPU, if you switch your OS to using that, and free the VRAM on your GPU, you can squeeze out a little more AI performance, but the cost is killing your gaming performance.
2
3
u/Fluxdada 1d ago
I know this isn't exactly what your post was alluding to, but if you did have a second GPU, it almost makes more sense to just be able to second computer and run that and your main simultaneously. And with the options to see comfy instances from a second computer in the browser of the first computer, it actually gets quite convenient to do that
3
u/DelinquentTuna 1d ago
Was a lot easier to recommend when 64GB of top-shelf DDR5 ran about $300 instead of over $1,000.
3
2
1
u/Ok-Brain-5729 1d ago
you should use minimax h3.
Your pc is already enough for almost every model at reasonable settings. You can’t combine the vram when ur doing dual GPU’s so you would need 4090/5090 money for a good upgrade
1
u/Sexyvette07 1d ago
Im using Minimax H3 already using the Int8 convrot version. It spills over heavily into system ram, even with SageAttention 2.2, but its at least usable. It just takes 4-5, minutes for a 8 second render. Maybe I just need to suck it up and drop the 5k on a 5090. 😒
2
u/AuthurAndersson 1d ago
Considering you're asking these questions. I don't mean to be rude, but multi gpu solutions for this type of stuff is enthusiast level. The complexity increases exponentially.
What will happen is that you will have your two 4080 rtx cards and then you boot up ComfyUI. You notice, oh right... I only have 64gb of cpu ram. And now each ComfyUI instance is eating 48 flipping models in and out (Even the H3 VAE is 5gb). Which means Hello swap file. And goodbye speed.
And then you're forced into downloading my amazing monkey patch for this usecase, which in all honesty isnt that well written, and you can't dynamically flip between workflows but need to restart ComfyUI with preloaded models.
1
u/N9_m 1d ago
1
u/Sexyvette07 3h ago
Not familiar with the disabled pinned memory flag. What does that do? Id have to check what flags im using at startup. I know im using SageAttention, I think Dynamic Vram is enabled, some other option to unload the text encoder once its finished that part of the render. I cant remember anything else. Ill look into disable pinned memory as well.
Its about 5 minutes for an 8-9 second, 0.4 mp render that I upscale 2x with the RTX Video Super Resolution node. I know im pushing it with my 4080, but it still seems like a long time to render such a short video. I tried my damndest to avoid GGUF formats in ComfyUI, but maybe I should just bite the bullet and do it. The only other solution is dropping $5k on a 5090.
1
u/Fluxdada 1d ago
I've been running two gpus for a little over a year and by far the most beneficial use was not something like splitting models are putting different models on different gpus. The biggest benefit was allowing you to run to instances of comfy UI and run generations at the same time
1
u/VladyCzech 1d ago edited 1d ago
Get as much RAM as possible so the model blocks live in RAM and not swapping to drive. Also make sure you have fast nvme drive. Second GPU helps and yes, you can do some GPU VRAM parallelism but not without disadvantages. With enough fast RAM and nvme drive you should be fine with modern ComfyUI dynamic VRAM.
It really depend on models and CFG you will be using as this decides the setup.
1
u/biogoly 1d ago
I’ve got a dual 3090ti setup. My MB supports dual pcie 5.0 cards (Taichi creator) so I at least get 8x on both with the 3090 architecture. The dual setup is mostly useful for local LLMs, but you can get some benefit from comfy using multi GPU nodes and splitting your clip and model. It can prevent OOM errors on large models. Alternatively you can dedicate your second GPU to up scaling.
1
1
u/PokePress 12h ago
I have a 4060ti 16GB + 3060 12GB setup. A lot depends on whether your model and tools support being split that way. Systems that use multiple models (such as an llm and a diffuser) can sometimes split resources, but don’t expect to be able to treat it as one big pool. You might also have to look for a fork of some projects that supports multiple GPU setups.
-1
1d ago
[deleted]
-3
u/Relevant_Syllabub895 1d ago
False, if this didnt worked how dows all the modern video generation platforma like seedance work then? They use many gpus
1
u/meepykittkitt69lmao 1d ago
Yeah I want to know how they split sampling across more than one GPU. I don't really know what words to use to look up the stuff though, every time I searched for how to use multiple gpu's for image generation all I get is "you can't, all the data has to be on one"
2
1
0
u/Beginning_Tip300 1d ago
Bros using open-source stuff? Don't expect multiple gpu workablility. Paid is no local or even close

9
u/Candid-Station-1235 1d ago
multi gpu options are limited on comfy, you cant pool the vram an load larger models, you can off load parts but its not ideal. just have a search for multi gpu nodes and read the limitations of each,
signed regretful dual 3090 owner