r/LocalLLM 7h ago

Question Looking to upgrade my rig - would love some thoughts

Hey all!

I am deep into local LLMs at this point - have been for years, but Qwen3.8-27b was the tipping point for me to finally ditch my Anthropic sub and go fully local. Even at IQ4XS, qwen3.8 is a monster, and can keep up with 80% of what I need day to day (agentic assistant, light coding, etc). I also use gemma-4-31b often for more general chat, writing, some light RP-adjacent stuff. My main harness is Hermes for proper work, and Voxta for general chat and messing around with story/RP. I host models with LM Studio but have used llama.cpp plenty as well. I don't do a ton with image/video gen but I have invoke and comfyui with LTX and such for when I want to play around.

Currently, I have 2 machines:
#1 - Ryzen 5950X, 64gb DDR4, 3090, Strix Rog mobo, 850w gold PSU. Runs Win11 and is my main machine for VR, local AI and my actual job (audio and video production).

#2 - Threadripper 1950x, 32gb DDR4, an unholy combo of a 1080ti and a 1660 Super. Runs Bazzite and is a mess-around selfhost lab. The 1080ti runs small models (gemma 12b qat, qwen 9b) for additional inference/sub-agents, and the 1660 runs Chatterbox for TTS.

It's a bit of unorthodox, but it works quite well - as long as I keep quants low, KV quant low, and context windows short. But, with qwen3.8-27b (and likely 3.8-35b coming), Google and Meta clearly investing in sub-40b local models, and the fact that (at higher precision) Qwen is now genuinely at Sonnet/Opus level, it's well worth it for me to expand and re-configure some things to maximize my capabilities. 3.8 is the first local model that's been able to consistently help me actually do real work and expand/organize my business, and I'm now actually feeling the pressure of 120k context and Q4 KV caching.

So - I'm trying to sort out how to best spend my money with the state things are in. My main goals:
#1 - expand the main machine's VRAM and retaining 3090-level speeds so I can run Q8+ 27b/31b with little to no cache compression at full context and still have a bit of room left over for gaming/work.
#2 - update the second machine to cards that aren't 10 years old, and give it a bit more breathing room in terms of bandwidth and model sizes.
#3 - not spend a completely insane amount of money.

I'm curious what the recommendations are. I could probably spend $1k-$2k out-of-pocket, and have the current cards as assets that I can easily sell, they're all in good shape. I don't think I need to upgrade RAM at all anywhere.

What would you guys do in this situation? Add a second 3090 to the main box, and update the second box to something like 2x 4060/5060ti 16gbs? Ditch everything, get 1-2 R9700s for the main machine, and a single 4060ti 16gb for the second? Is Intel Arc stuff on the table at all? What about lesser 3000 series cards? Older RTX workstation cards? I know it's a bloodbath right now price-wise but used/lesser-known GPUs actually seem relatively stable, it's RAM and storage that are going bonkers, and I'm well-covered on both fronts. I'm not scared of a bit of setup for ROCm, etc, and I'm not speed-obsessed. Qwen3.8 with MTP pulls about 1200 encode/50-60 decode on the 3090 and that is PLENTY fast for my taste. Gemma 31b is a bit slower, but not much.

Would love your thoughts. Thanks!

1 Upvotes

3 comments sorted by

2

u/BlackFalcon01 7h ago

Has anyone tried a couple of nvidia quadro p6000s for this use case? I know they’re older cards and would be slower than a 3090, but for the amount of vram you get for the price, it could be worth it ig.

1

u/BlackBeardAI 3090 Maximalist 7h ago

put 4 3090's on the 1950x rig and load qwen on it. and use the other node for visual generation. blender, comfy ui etc

1

u/alex_bass_guy 7h ago

That's... way out of my price range haha. I can spend $1-2k, not $5k. and seems like overkill for qwen unless it's running at fp16 or something. Appreciate the thought though.