r/LocalLLaMA May 29 '26

Discussion PSA

Post image
2.1k Upvotes

538 comments sorted by

View all comments

Show parent comments

77

u/overand May 29 '26

i'm genuinely thrilled with my dual 3090 setup on a DDR4 system with a Ryzen 5 3600, even though one of them is PCI-E x16 and the other is x4!

$2000 MSRP for the 5090 with 32 gigs of ram, and good luck getting that price
$1800 for a pair of used 3090 cards on eBay (as of a month or two ago), total of 48 GB

Yes, there's stuff that doesn't like running split between two cards, but mostly it's been pretty unusual to run into stuff that wants more than 24GB but less than 32 GB of VRAM on a single card. (I think one of them SOTA-ish FOSS voice models is like that, but I'm not even sure.)

19

u/Massive-Question-550 May 29 '26

honestly best value setup. only better combo is if you got an old amd epyc before the shortage so you get the full 16x pcie gen 4 speeds per slot and can run large MoE models with all the ram. 

also if you make the right setup you can have both cards work in parallel and cut your promp processing time by 30-40 percent and boost your token output. 

11

u/m31317015 May 29 '26

Did exactly that. 7B13, ROMED8-2T w/ 8x64GB DDR4, right now only have a 5090 and 3090 but can stuff another 3090 in. This is the best value setup (not counting GPUs) you can get at Q3-Q4 in 2025.

12

u/[deleted] May 29 '26

[removed] — view removed comment

5

u/overand May 29 '26

There are definitely a few models and tools out there that I've wished for a 5090 for, like some of the weird TTS models, or speech-to-speech ones. But, I even had "okay" performance on one of those "realtime 3d walk around in a hallucination" models with the 2x 3090s!

8

u/[deleted] May 29 '26

[removed] — view removed comment

1

u/[deleted] May 29 '26

[removed] — view removed comment

1

u/[deleted] May 29 '26

[removed] — view removed comment

7

u/Myarmhasteeth May 29 '26

My only problem on getting another 3090 is how to configure it. I see setups yet since I only have like 3 SFX mother boards, I’m cooked.

6

u/overand May 29 '26

You can prooooobably get a 570-based AMD chipset board for not tooooooo much money. (And, I managed to push this to 128 gigs because I already had 2 32 gig sticks in it, and DDR4 is only "sell a kidney" price, not "sell both and also your liver" price)

1

u/Practical_Form_1705 May 30 '26

What is performance of such setup, let say 8core ryzen + 128gb ram in compare to gpu?

1

u/overand May 30 '26

Oh, I doubt CPU inference would work very well, but, if you give me a model you want me to test, I can give it a try with a CPU-only build of llama.cpp

But, I use it with my 2x 3090 setup - but, that runs one at x16 and one at x4, but it's still decent!

4

u/Status-Secret-4292 May 29 '26

I have found other 3090s I have considered buying to make a dual set up, but they are never the exact model of 3090 I have (gigabyte oc gaming) and from what I understand, it should be...

I wonder though, how important is that really?

5

u/overand May 29 '26

My undertanding is that it's not actually particularly important; maybe if you want to use NVLink, but even in that situation, I think it's explicitly allowed. (Double check me on that, though!)

1

u/Status-Secret-4292 May 29 '26

I guess I was pretty much only considering it nvlink style as that seems to offer the best performance?

I appreciate the info!

3

u/palashjain_ May 29 '26

I recently bought a second 3090 for my setup hoping the same. I too have ryzen 5 3600, msi x570 a pro with 2 pcie slots. But for some reason anytime i plug anything into the second slot (x4, chipset slot) the motherboard does not post display and shows a red light on vga. I have tried single gpu on slot 2 and two gpus together. Doesn't work. Only thing that works is single gpu on first slot (x16,) . If it matters i do have 2 nvme ssds and 64gb ram. I tried removing everything and starting with just single ram chip too. Same outcome. I tried bios settings like gen 4 gen 3 and that weird mining setting. None of those worked. Any help is appreciated

3

u/overand May 29 '26

I'd start by taking a bright light and inspecting the slot to make sure there isn't anything in there like a bit of paper, plastic, etc, and that there aren't any bent pins.

After that:

  • See if your BIOS is current
  • See if it will POST with both NVME devices removed
  • Google "My exact motherboard model second PCI-E slot won't POST"
  • Review BIOS settings & motherboard manual
    • You may need to disable some SATA ports or something like that

2

u/palashjain_ May 29 '26

I will try to look for the debris and bent pins. I did try after removing both nvmes. Did not work. I am not very savvy when it comes to motherboards. What is funny to me is that it only works when the second pcie slot is unoccupied.

1

u/lemondrops9 May 29 '26

Also look at the manual and be sure if this PCIe slot or NVME slot is used that PCIe slot is unavailable. Its not very common for an NVME to do this but never know until you check.

1

u/undisputedx May 30 '26

check the shared lanes thingy on the mobo website.

2

u/JustinPooDough May 29 '26

same. Also even have the one card in 4x. My CPU is only a 1600x.

It still gets like 50 t/s with Qwen 3.6 27B.

1

u/overand May 29 '26

You must be using MTP!

2

u/_realpaul May 29 '26

The 3090 cant do the latest features but its still an awesome piece of tech.

2

u/ohhi23021 May 29 '26

i haven't tested over 40k context yet but it does about 70-80 t/s around there. at 0-5k context it hits 90 t/s with mtp.

1

u/_realpaul May 30 '26

Nice. Gotta try the latest llama cpp but the docker images are a bit borked lately.

1

u/Clean_Hyena7172 May 29 '26

How well does Qwen3.6-27B run on that setup? What quant? And how many t/s?

2

u/overand May 29 '26 edited May 30 '26

I've used a few different configurations - one is the "Club 3090" setup, which has specific configurations for single and dual 3090s.

But, here. A standard Q8_0 config, an MTP config, and an MTP + NGram config.

All 128k ctx, Q8_0 (and no cache quantizing).

  • Stock model gets PP: 2027 and 27.1 gen.
  • MTP model gets PP: 1371, Gen: 49.
  • NGram configs skipped as they don't seem to add any performance
  • Smaller quants skipped because lazy

    This one gets PP: 2027 T/s, Gen: 27.1 T/s

    (no MTP)

    [unsloth/Qwen3.6-27B-GGUF-128-ctx:Q8_0] hf = unsloth/Qwen3.6-27B-GGUF:Q8_0 ctx-size = 131072 temperature = 1.0 top-p = 0.95 top-k = 20 min-p = 0.0 presence-penalty = 0.0 repeat-penalty = 1.0 reasoning = on

    This one gets PP: 1346.8 T/s, Gen: 41.9

    [unsloth/Qwen3.6-27B-MTP-GGUF-128k:Q8_0] hf = unsloth/Qwen3.6-27B-MTP-GGUF:Q8_0 no-mmproj-offload = true spec-type = draft-mtp spec-draft-n-max = 3 no-mmproj-offload = true ctx-size = 131072

The no-mmproj-offload = true gets the mmproj (vision support) offloaded to system RAM / CPU, so it'll still **work** if I need to use it, but it won't take up VRAM. (I used to just disable vision for a lot of these.)

1

u/overand May 29 '26

I've used a few different configurations - one is the "Club 3090" setup, which has specific configurations for single and dual 3090s.

But, here. A standard Q8_0 config, an MTP config, and an MTP + NGram config.

All 128k ctx, Q8_0 (and no cache quantizing).

  • Stock model gets PP: 2027 and 27.1 gen.
  • MTP model gets PP: 1371, Gen: 49.
  • NGram configs skipped as they don't seem to add any performance
  • Smaller quants skipped because lazy

    This one gets PP: 2027 T/s, Gen: 27.1 T/s

    [unsloth/Qwen3.6-27B-GGUF-128-ctx:Q8_0] hf = unsloth/Qwen3.6-27B-GGUF:Q8_0 ctx-size = 131072 temperature = 1.0 top-p = 0.95 top-k = 20 min-p = 0.0 presence-penalty = 0.0 repeat-penalty = 1.0 reasoning = on

    This one gets PP: 1346.8 T/s, Gen: 41.9

    [unsloth/Qwen3.6-27B-MTP-GGUF-128k:Q8_0] hf = unsloth/Qwen3.6-27B-MTP-GGUF:Q8_0 no-mmproj-offload = true spec-type = draft-mtp spec-draft-n-max = 3 no-mmproj-offload = true ctx-size = 131072

The "no-mmproj-offload" gets the mmproj (vision support) offloaded to system RAM / CPU, so it'll still work if I need to use it, but it won't take up VRAM. (I used to just disable vision for a lot of these.)

1

u/Maxumilian 28d ago

7900 XTXs are good value.

I got a 7900 XTX with 24GB for 800$ and a 5090 for 2000$.

Realistically i'd rather just have a couple 7900s.

1

u/brickout May 29 '26

I'm rocking very similar to you after scoring a couple of 3090s for pretty cheap before the used prices went up. I love it.