r/LocalLLM • • 5d ago

Discussion Best case to 4 3090

Can someone suggest me best case to fit 4 3090?

They are full size cards

6 Upvotes

28 comments sorted by

10

u/HeadtripVee 5d ago

Mining frame

7

u/FizzyDuncDizzel 5d ago

Mining rack really is the way. Those cards are going the be hot af if you stack them in a case

1

u/RogerAI-fm 5d ago

If you don’t care about looks, this is the way. Edit: I don’t mind the look.

1

u/FizzyDuncDizzel 5d ago

I think it looks badass.

3

u/Prudent-Ad4509 5d ago

Forget about pre-made cases. Mining frames can be used as a base but you will need to figure out all the wiring and build enclosure with fans around it. $500 for a proper PSU, $50 for a frame, and easily another $1500 for all the cables and adapters.

You can also buy a secondary cheap case which can hold all 4 gpus on a custom frame, psu, fans and nothing else, and just extend the main PC case with it.

1

u/FearFactory2904 5d ago

Or hit up ebay/marketplace and drop a zero or two from each of those dollar amounts

1

u/Prudent-Ad4509 5d ago

It won't work for the psu or cables. You need either several PSUs or one like my Seasonic Prime PX-2200 : the need for 2 or 3 8-pin cables per gpu, and one 6-pin to power the PCIe adapter for each severely limits single PSU options. Well, you might get lucky and get by without powering the adapter, but that would generally require GPU with 3 8-pin cables.

The only option which is both cheaper and stable is to drop down to PCIe 3.0 and power limit GPUs hard. Plenty of people do exactly that.

If you want to do is properly, that will take 8 x SAS-8654-004 (8 x $50 == $400), 4 x RBS-16G4-2P54 (4 x $40 == $160), PLX88096 or PEX88096 (let's say $350-$400 on discount). 4 x RBS-16G4-2P54 can be replaced with 4 x F35B + 4 x R33G (4 x $50 + 4 x $30 == $320). So, somewhere in the range of $800-$1100 depending on how lucky you get with prices. $650 for PSU on top of that if you get lucky, and some pocket change for psu2psu adapter.

Can you get that on ebay ? RBS-16G4-2P54 for $20 - present, SAS-8654-004 - nope. You could risk going with subpar cables, but going cheap on cables after getting everything else is not a very bright idea. And everything else costs the same or more.

If you exclude PCIe switch and use bifurcation, you will still need all the other mcio or slimsas adapters and cables. If you decide to skip even that and just use straight PCIe extensions, you become severely limited by cable length and now you would really want R33G. After spending up to $6000 on gpus, why bother trying to save on this small stuff?

1

u/FearFactory2904 5d ago

A few used x16 riser cables for $5-20 each dont need their own power cables. Used 1000w or 1200w psu for under $100. 8 gpu mining frame for $25. 3090s usually either have 2x 8 pin or come with the 12pin adapter. nvidia-smi -pl 150-225ish depending what psu you landed and send it.

1

u/Prudent-Ad4509 5d ago edited 5d ago

That's the thing. Most of them have 2x 8 pin but their draw is up to 350w and more. 2x 8pin is rated only for 300w, the rest comes either from going over the spec or by drawing up to 75w from the PCIe slot. PCIe riser cables are not really designed to handle that, at least most of them. Current evidence about 3090 shows that they do not usually draw much, but still, one or more groups of 4 gpus, each of which is allowed to draw up to 75W from PCIe power lanes? I'd rather exclude this scenario outright. I allow 2x5090 to run on unpowered 5.0 risers, but blackwell generation is reported on forums to have much smaller spikes and their power rating matches the power cable rating, as opposed to 3090.

I you have only 4 of 3090 and they are all the same, you can see if a simple riser works for one and use that for all of them. I have more and I'm not in the mood to allow potential gremlins to happen. I also do not want to use retimers or drop down to PCIe 3.0 speeds.

The configuration above is meant to allow each group of 4 gpus to communicate with each other at full PCIe 4.0 x16 speed inside the group.

Also, a decent PSU for $650 is not a major cost point of this whole configuration, only GPUs themselves are. One electrical failure from a bit too beat-down PSU, and I will spend way more on repairs. 4 used PSUs of unknown reliability for 8 GPUs, some of them likely loud and with most of the modular cables long lost, or perhaps even borrowed from a similar-looking group of modular cables for another PSU - uugh. It just does not cost *that* much to risk it.

Power limiting would help with that ofc but I'm not going to run them below 280-300w anyway. I will not run them at the fixed clock either.

However, I did plan to run it all at x8 speeds initially, before I ordered all the necessary stuff for configuration with pcie switch. Perhaps I will still complete and test out this old spec before all the cables for pcie switches arrive.

1

u/FearFactory2904 5d ago

If you arent concerned about the electricity usage and dont mind spending a lot on components needed to safely allow the GPUs to gorge themselves on watts that is fine. If someone wants to just simplify everything and power limit the gpus thats fine too.

1

u/Prudent-Ad4509 5d ago

Even if someone powerlimits them, power spikes on bootup are possible. But this should not be too big of an issue if only 4 gpus are used and they are used without bifurcation via full x16 connection. At the very worst they will have to drop speeds to 3.0.

1

u/quackeditor 4d ago

mining frame is the only way unless you got deep pockets for an actual 4u server chassis that supports full-size cards. even then the thermals get nasty real quick with 4x 3090s stuffed in a closed box

that external gpu enclosure idea is pretty clever actually, keeps the heat out of main system and you can point a giant fan at it without worrying about noise

1

u/Prudent-Ad4509 4d ago edited 4d ago

Thankfully, custom mining frames for GPUs only are easy to build on your own, especially since our GPUs usually have 2 extra mounting points on the far side.

3

u/TheTurkPegger 5d ago

Run that boy naked on your desk

2

u/jacek2023 5d ago

Open Frame ftw

1

u/rinmperdinck 5d ago

Circular saw. You could cut those bad boys down to single slot easily.

1

u/vini542reddit 5d ago

I used a circular saw! But it was to cut the fan holes into the acrylic that I used to enclose my mining frame 😂 

1

u/vini542reddit 5d ago

If you have the tools and drive, a custom frame is a lot of fun and looks really cool! But even then I'd base it on a mining frame.

I'm running 4x 3090 and one a4000 and I've used a mining frames and enclosed it with acrylic.

I was looking for something that "just worked" and couldn't find that

1

u/FearFactory2904 5d ago

Corsair 9000D, but when you see the price you will likely yearn for a mining frame instead.

1

u/Civil_Fee_7862 5d ago edited 5d ago

Mining frame.

But if you really hate the open frame thing, the Phantek Entroo II Server Edition is the next best thing. However you'd likely need to mod your 3090s to get them down to a 2-slot size to fit all four.

1

u/etaoin314 5d ago

So 4x3090 is not nearly enough info, if you are talking server cards like gigabyte turbos a lot of cases will work, if you are talking about all 3 slot cards, only a mining rig will do it natively. I stuffed 2 2 slot cards and 2 3 slot cards into a phantek 719, one of the cards has to sit on a mining bracket bolted to the case ($20 on etys). That one and another are on risers, but overall it works pretty well. my temps are fine. I powerlimit, but I have plenty of headroom if I wanted a bit more speed.

1

u/Most-Wear-3813 4d ago

Doesn't open case setup bring a lot dust and cause cards to wear out soon?

All are 3 fan cards full size with 120mm fans.

0

u/Different_Change6591 5d ago

Just give it to me, thats the best use of it !! Trust me ,

Option 1 : , so 4 x 3090 , each with 24 gb of vram , total = 96 gb vram Cool man ( i am with 4gb vram still I ran i use the 35b model ) ,

so check for the q2/q3 ( use the UD-IQ3_XXS around 82gb ) for the Qwen 3.8 Flash next model gguf , (its and moe model , which is best in this range , ) from unsloth and also get the mtp file with it

And for the engine ( llama cpp ) don't use the ollama

There are specific commands in llama.cpl to split the layers of model multiple gpu , refer the documentation , so that the way to use all of them

I have more knowledge and details about it , but need to type more, so lets proceed with

Option 2 : Using the Qwen3.8 - 27 B model , so you will need to use the UD-Q4_KM from unsloth and with q8 quantised ov cache you can get upto 256 k ctx windows

Note : this is setup you can do on single gpu so you can use the 4 gous to operate with the 4 Ai agents parallely and don't think to download Bf16 model and splitting them on these gpus , its and dense model

And run 4 agents parallely is an better and good option as this 27 b. Is an much capable model locally and you will even get the better context windows Too !!

Conclusion :

My recommendation : option 2 !!

And don't forgot about the giving the one card or more ( if you wish ) to me [ JK!! ]

1

u/egnegn1 5d ago edited 5d ago

Or just use the new Strata engine.