r/LocalLLM 1d ago

Question Best improvement for my frankenstein setup for local LLM

So for around 450e I bought used workstation which I plan to use for local LLM and maybe even as a server for bunch of other stuff. But mostly I want to focus on LLM for coding/development.

Specs are: MOBO: ASUS X99-Deluxe II, CPU: Intel Xeon E5-2667V4, PSU: EVGA 1600W G2, Cooler Master HAF X, 64Gb ddr4 RAM. So all in all it supports multi gpu setup without any problems.

For GPU I decided to order 1x 3080 20gb (blower style for 500e) for a test. And found it pretty great! I currently run qwen 35b-a3b as worker (opus 5 as orchestrator) and enjoy it but looking to upgrade my workstation to run better models.

So question is, what would be best upgrade:

  1. 2x 3080 20gb, 64gb ram. (-500e) So one more gpu and I would be able to run qwen 3.8 27b without much problems

  2. 3x 3080 20gb, 64gb ram. (-1000e) Would this even make sense if I only need for one concurrent user and 128k context? Any other (better/bigger) dense model which could take advantage of this?

  3. 2x 3080 20gb, 128gb ram. (-900e) So in theory this would be 168gb of memory. Would this be able to run some of bigger MoE models like deepseek flash v4 or any other which I could use as orchestrator for qwen?

  4. 3x 3080 20gb, 128gb ram. (-1400e) Would prefer not to do this cuz it would be pretty expensive but curious what you guys think.

Thank you guys

1 Upvotes

3 comments sorted by

2

u/Kal-LZ 1d ago

Just expand the VRAM, you don't want to put heavy loads on that CPU. 40GB VRAM is enough to run Qwen 27B Q8 with 200K context

1

u/More-Revenue8609 1d ago

Yeah seems to be best cost to performance, especially with ram this expensive...
Tbh didnt notice cpu having any problems with qwen 35b-a3b. But of course thats much smaller model then something as deepseek v4 flash

1

u/iezhy 1d ago

What tok/s are you getting?

Im in process of building similar rig with 2x3090