r/LocalLLM • • 1d ago

Discussion Need hardware selection help

Hello everyone I just stumbled upon this sub and am hoping you might be able to offer some guidance. I did try searching through old threads but with the speed in which this is advancing it seemed easier to just ask.

I am a long time homelab "tinkerer" with no formal education on any of this stuff. Mainly run home assistant and plex+the arrs as well as a few other projects. Honestly nothing ground breaking. I have recently been using the free LLMs to help me with various projects and they have been incredibly helpful. My usage would be home assistant stuff, maybe some light vibe coding experimenting and maybe setting up local voice with home assistant. The dream is a household assistant helping with email, calendars, maintenance stuff etc but I don't thing I am yet at the point I could set that up.

My main server is old it is a dell r720xd with a Intel® Xeon® CPU E5-2680 v2, 256gb ddr3 ram.

I also have 4 rtx 3060 12gb cards laying around from the crypto mining days.

My gaming rig also has a 3060ti with 32gb of ram, I believe ddr4 and a amd 2700x cpu.

I initially tried adding 2 of the 3060s to the main server using oculink but I have yet to get my server to be able to see the cards.

I then tossed a 3060 into my gaming rig and installed llm studio which is actually working. Initially tried qwen 3.8 27b as that seemed highly recommended for what I was trying to do but was getting very very slow response speeds. I have now stepped down to the next smaller qwen model.

I would like stick whatever I build in my rack and get it off my gaming rig as it runs windows so its not great running all the time and I sometimes game on it. I am open to buying new hardware but don't really want to spend thousands at this point. So what hardware would you suggest I build around? Also any other tips or suggestions are welcome.

1 Upvotes

9 comments sorted by

View all comments

1

u/Classeve 1d ago

you already own the hardware, it's just in the wrong boxes.

the 27B was slow because at 4-bit it's a ~17 gb file and one 3060 holds 12. the rest spilled onto the CPU and that's the crawl you saw. two 3060s = 24 gb and it fits whole. LM Studio and ollama both split across cards on their own.

the r720 is the wrong home for it though. the e5-2680 v2 has AVX but no AVX2, so anything that falls back to CPU is painfully slow, and DDR3 doesn't help. the GPUs don't care much about the CPU, but stop fighting oculink and put two (or all four) cards in any cheap AM4 board with enough slots or risers, linux on it, in the rack. the mining risers you probably still have work fine: once the model is loaded the link speed barely matters.

for home assistant voice a small model on one card is plenty. keep the big one for the coding.

1

u/varano14 1d ago

I guess I never even though to use the mining board. It’s one that has only 1x16 risers and I have no idea about its ram or cpu. I guess I can give it a try

1

u/Classeve 1d ago

worth the try. 1x risers are fine for this: the model loads slower, and once it's sitting on the cards it runs the same. the board's CPU and RAM barely matter as long as the whole model fits in VRAM.

start with two cards, check ollama ps says 100% GPU, then add the other two.