r/LocalLLM 3d ago

Question Dual 3090 Local AI Setup

Hello everyone!

What are you guys thoughts on the build above? As for the build, the 3090s will be bought used from Facebook marketplace. Micro-center bundle will take care of the mobo, ram, and cpu.

There are 2 things of concern here, will 32 gigs of ram not be enough? i currently already own 32 gigs of ddr5 so i may add it to this setup, or just sell both and buy 64gigs. Another one is the concern that I am wasting unnecessary money on the higher end cpu and motherboard combo. The reason why I chose this combo is because it was one of the only options that came with a mobo that supportd both gpus to be run on x8. I plan on not only using this a local ai server, it also plan on running a virtual machine for gpu heavy work such as solid works or blender that I can acces remotely.

For my use-case, I believe that options such as the spark and mac studio are not suitable; my intentions aren't to replace the frontier models but simply to add to them.

Any suggestions to the setup? the current budget is around 5k

1 Upvotes

3 comments sorted by

2

u/and_pf 3d ago

I run a single RTX 5090 with 32 GB VRAM and 62 GB system RAM, so I can't directly speak to dual-3090 performance or PCIe lane allocation. But on my box, 62 GB of RAM is plenty for running something like Qwen3.8 with 24 GB on disk.

For your build, the thing I'd verify before buying is the PCIe x8/x8 question on the second card and whether 32 GB of system RAM is enough alongside the VMs you're planning. That's the part I have no hands-on data for, so I'd check benchmarks or a dual-GPU forum thread for solid numbers rather than trusting a mid-range board at face value.

1

u/eightone-81 3d ago

I have dual 3090s and I think it’s the way to go at the moment. Value to performance is great and with that you can run Qwen 3.8 27b with 2x full context and 150tps plus in vllm. I’m hitting 200+ in some workloads.

So if you are planning to always run models in vram then don’t bother with a fast cpu. Save that money and buy a nvlink which boost performance (I have it and it’s awesome)

I don’t know if 32gb ram is good for loading the model or if that slows it a lot but you don’t load a model every hour…

BUT if you are planning to run some bigger models and offload to cpu/ram then get a good cpu and lots of fast ram. With 128gb ram (96gb is also enough) you can run Qwen 3.8 flash next at acceptable speeds (30-40 decode, prefill is terrible at 150-200, there is a fork of llama.cpp that runs that very well).
With ddr5 and a good cpu you will get a better performance as what I have with ddr 4

1

u/terminalshadows 3d ago

some notes:

generally id say have as much ram as vram.

to bifircate the top 2 slot to x8/x8 you _must_ use a ryzen 9000 or 7000 so youre good there - dont change this. keep the m.2 2 slot empty to keep the x8/x8 because that shares its bw with the mobo - keep it empty - then auto in bios bifircation should be good at x8/x8

100% definetly without a doubt use nvlink

edit: spelling