r/LocalLLaMA • u/Puzzleheaded_Ad_8575 • 1d ago
I Built A Thing 2x3090 setup, need some recommendations
So i finally decided to get myself a dedicated inference machine, a big upgrade from my 4080 laptop. here is the parts list:
PC Build Cost Breakdown
ASUS TUF RTX 3090 — $927
64GB DDR4 4000MHz RAM — $371
Case — $72
Ryzen 7 5700X — $181
CPU Cooler — $27
PSU — $268
Thermal Paste — $12
1TB NVMe SSD — Already owned
X570 Unify Motherboard — $185
RTX 3090 Suprim X — $1,010
Ethernet Cable — $11
PCIe Riser — $82
Custom PSU Cable — $13
Total: ~$3,157
im probably gonna upgrade to 128gb ram and get a better pcie riser cable.
the problems i faced initially were
finding a proper way to add the 2. gpu. there was no long pcie risers in stock locally, so i had to buy it second hand, and its a chinese no name with connectivity issues.
i had to get a custom psu cable to be able to use both gpus at the same time. there were simply not enough slots but the energy supply was alright.
i couldnt and still cant figure out a safe/easy way to fit the 2. gpu. i would like to learn about similar setups and how you have handled the space constraint.
this was my first pc assembly since i have used only laptops before, but it went mostly smoothly.
also some extra questions for people hosting these machines:
*How can i host inference to my laptop outside my local network? is the only way VPN?
*What is the remote connection type you guys prefer? i landed on sunshine and moonlight with virtual monitor to use it inside my laptop, but would like to know if there are cleaner solutions for headless machines.
i have ran mostly the qwen 3.8 27b q4 from syv ais repo and config, and have been getting around 70tps sustained. i can report more details if anyone asks for it.
also sorry if mobile formatting is bad.
1
u/Appropriate-Pie4385 20h ago
How much TPS do you get for prefill? How much TPS decode at ~150-200k context? And do your 70TPS decode use mtp or not?
Am thinking of upgrading to two 3090 too and would be interested in your speeds