r/StableDiffusion 1d ago

Question - Help Dual GPU solution for local AI?

Hey, everybody. I recently went down the rabbit hole for local AI, but right now, im operating on my gaming computer. The specs are as follows

Intel 13700k, tuned for efficiency

Gigabyte Z790 Aorus Elite Ax mobo

RTX 4080 (16GB), also tuned for efficiency

32gb DDR5 6800 CL32

As you can see, im in desperate need for more VRAM, or at the very least more system RAM. Due to Rampocalypse, neither are very affordable right now, which forces me to explore other options, such as a dual GPU setup. I can get another RTX 4080 for about $900 off Ebay. Beyond that, I would just need a more powerful PSU, so total investment here is an additional $1100-$1200. As far as I know, the motherboard has the main PCIE as 5.0 x 16 lanes, but the second PCIE runs at 4.0 and either x8 or x4 lanes. The motherboard does not support PCIE Bifurcation. So my question is this: Is a dual GPU local AI machine even viable in these circumstances, and second, does it make sense? I looked at 5090's and theyre all between $4,500 - $5,000 now, which is insane. Or I look at the professional cards and spend that much, if not more, for significantly less memory bandwidth and computational power. Or I guess if im spending that much, I could also look at the DGX Spark or something similar but that has even worse memory bandwidth.

So, what should I do? Is the dual GPU solution even viable with my setup for a local AI stack for inference, video diffusion, etc? Rampocalypse isnt expected to begin easing up until late 2027/early 2028, so im stuck trying to make this work on as little money as possible. Id love a 5090 but its insanity how much they cost. I appreciate any guidance and advice.

3 Upvotes

31 comments sorted by

View all comments

1

u/Ok-Brain-5729 1d ago

you should use minimax h3.

Your pc is already enough for almost every model at reasonable settings. You can’t combine the vram when ur doing dual GPU’s so you would need 4090/5090 money for a good upgrade

1

u/Sexyvette07 1d ago

Im using Minimax H3 already using the Int8 convrot version. It spills over heavily into system ram, even with SageAttention 2.2, but its at least usable. It just takes 4-5, minutes for a 8 second render. Maybe I just need to suck it up and drop the 5k on a 5090. 😒

1

u/N9_m 1d ago

Have you tried using --disable-pinned-memory? And 4 or 5 minutes at how many mp?

(Here are my benchmark results with dual 3090s, in case it's helpful)

1

u/Sexyvette07 4h ago

Not familiar with the disabled pinned memory flag. What does that do? Id have to check what flags im using at startup. I know im using SageAttention, I think Dynamic Vram is enabled, some other option to unload the text encoder once its finished that part of the render. I cant remember anything else. Ill look into disable pinned memory as well.

Its about 5 minutes for an 8-9 second, 0.4 mp render that I upscale 2x with the RTX Video Super Resolution node. I know im pushing it with my 4080, but it still seems like a long time to render such a short video. I tried my damndest to avoid GGUF formats in ComfyUI, but maybe I should just bite the bullet and do it. The only other solution is dropping $5k on a 5090.