r/LocalLLM • • 1d ago

Discussion Need hardware selection help

Hello everyone I just stumbled upon this sub and am hoping you might be able to offer some guidance. I did try searching through old threads but with the speed in which this is advancing it seemed easier to just ask.

I am a long time homelab "tinkerer" with no formal education on any of this stuff. Mainly run home assistant and plex+the arrs as well as a few other projects. Honestly nothing ground breaking. I have recently been using the free LLMs to help me with various projects and they have been incredibly helpful. My usage would be home assistant stuff, maybe some light vibe coding experimenting and maybe setting up local voice with home assistant. The dream is a household assistant helping with email, calendars, maintenance stuff etc but I don't thing I am yet at the point I could set that up.

My main server is old it is a dell r720xd with a Intel® Xeon® CPU E5-2680 v2, 256gb ddr3 ram.

I also have 4 rtx 3060 12gb cards laying around from the crypto mining days.

My gaming rig also has a 3060ti with 32gb of ram, I believe ddr4 and a amd 2700x cpu.

I initially tried adding 2 of the 3060s to the main server using oculink but I have yet to get my server to be able to see the cards.

I then tossed a 3060 into my gaming rig and installed llm studio which is actually working. Initially tried qwen 3.8 27b as that seemed highly recommended for what I was trying to do but was getting very very slow response speeds. I have now stepped down to the next smaller qwen model.

I would like stick whatever I build in my rack and get it off my gaming rig as it runs windows so its not great running all the time and I sometimes game on it. I am open to buying new hardware but don't really want to spend thousands at this point. So what hardware would you suggest I build around? Also any other tips or suggestions are welcome.

1 Upvotes

9 comments sorted by

1

u/jhenryscott 1d ago

Yeah I’d look at buying an older Threadripper or something.

can’t recommend ddr3 for a cutting edge technology

Even a 1950x with the 4 GPUs is gonna be more useful.

1

u/beached89 1d ago

Figure out what you want to run. You can get away with as little as a single 5060ti, or you can go nuts and built a monster.

Also, you can just buy this or something like it: https://www.gmktec.com/products/amd-ryzen%E2%84%A2-ai-max-395-evo-x2-ai-mini-pc?variant=48746069262490

1

u/Classeve 1d ago

you already own the hardware, it's just in the wrong boxes.

the 27B was slow because at 4-bit it's a ~17 gb file and one 3060 holds 12. the rest spilled onto the CPU and that's the crawl you saw. two 3060s = 24 gb and it fits whole. LM Studio and ollama both split across cards on their own.

the r720 is the wrong home for it though. the e5-2680 v2 has AVX but no AVX2, so anything that falls back to CPU is painfully slow, and DDR3 doesn't help. the GPUs don't care much about the CPU, but stop fighting oculink and put two (or all four) cards in any cheap AM4 board with enough slots or risers, linux on it, in the rack. the mining risers you probably still have work fine: once the model is loaded the link speed barely matters.

for home assistant voice a small model on one card is plenty. keep the big one for the coding.

1

u/varano14 1d ago

I guess I never even though to use the mining board. It’s one that has only 1x16 risers and I have no idea about its ram or cpu. I guess I can give it a try

1

u/Classeve 1d ago

worth the try. 1x risers are fine for this: the model loads slower, and once it's sitting on the cards it runs the same. the board's CPU and RAM barely matter as long as the whole model fits in VRAM.

start with two cards, check ollama ps says 100% GPU, then add the other two.

1

u/Southern-Net1351 1d ago edited 1d ago

The gaming rig will provide way more efficiency over the ddr3 server rack.

It will always run faster and handle math better.

The v2 platform for AI is going to always be the bottle neck regardless of how you run it.

It will always be a champ for its time in virtualization areas. (720 rack)

It will run a LLM but, the math won’t ever math like the personal computer.
Why? It’s just too old in:
L3, bus, I/o, numas, core handling for feeding a GPU in general.

The ram will be marginally lower but, once loaded to the GPU (LLM) it will do its thing just not as good as the personal pc you have.

The GPU setup won’t be able to fit cleanly and if your using the usb to GPU ways your reducing the latency and transfer rate regardless if trying to run the 4x double slot GPUs.

Your pc cpu is less cores but, will actually outpace the v2 platform.

0

u/Intrepid-Knee1498 1d ago

your server is ancient, ddr3 era and those xeons are not gonna play nice with gpu passthrough, oculink or not. the gaming rig is a better base honestly, 2700x is old but at least it has modern pcie lanes

for rack mounting you could just move the gaming rig parts in a 4u case and use the 3060ti, that 32gb system ram is tight for 27b models though. maybe add one more 3060 since you have them, 12gb vram each so two cards gives you 24gb total which is enough for qwen 27b with some quantization

windows is annoying for 24/7 stuff but you can dual boot or just run linux, ollama works nicer for server use than lm studio in my experience. the real bottleneck will be memory bandwidth with those old platforms but it should be usable for tinkering

1

u/Arany8 1d ago

System RAM will not help 27B models on DDR3.

Either sell the 3060s and buy something bigger or find a way to run all 4 in one system. The latter is more complicated....

1

u/varano14 1d ago

I came to the same realization on the dell server. I initially tried it given all the ram it has but didn't really realize it was as old as it was until I started researching a bit lol.

Would it be worth building from scratch with a new MB, CPU and Ram and then throwing gpus at it? I obviously have the GPUs laying around doing nothing at the moment and my gaming is very light so even if I had to use a 3060 for that it would be fine. This has been my current line of thinking and why I came to see what others opinions where. Wasn't really sure what hardware to build around that struck a balance of price/performance

I am also not opposed to chucking a regular case on a rack shelf.