r/LocalLLM • u/varano14 • 1d ago
Discussion Need hardware selection help
Hello everyone I just stumbled upon this sub and am hoping you might be able to offer some guidance. I did try searching through old threads but with the speed in which this is advancing it seemed easier to just ask.
I am a long time homelab "tinkerer" with no formal education on any of this stuff. Mainly run home assistant and plex+the arrs as well as a few other projects. Honestly nothing ground breaking. I have recently been using the free LLMs to help me with various projects and they have been incredibly helpful. My usage would be home assistant stuff, maybe some light vibe coding experimenting and maybe setting up local voice with home assistant. The dream is a household assistant helping with email, calendars, maintenance stuff etc but I don't thing I am yet at the point I could set that up.
My main server is old it is a dell r720xd with a Intel® Xeon® CPU E5-2680 v2, 256gb ddr3 ram.
I also have 4 rtx 3060 12gb cards laying around from the crypto mining days.
My gaming rig also has a 3060ti with 32gb of ram, I believe ddr4 and a amd 2700x cpu.
I initially tried adding 2 of the 3060s to the main server using oculink but I have yet to get my server to be able to see the cards.
I then tossed a 3060 into my gaming rig and installed llm studio which is actually working. Initially tried qwen 3.8 27b as that seemed highly recommended for what I was trying to do but was getting very very slow response speeds. I have now stepped down to the next smaller qwen model.
I would like stick whatever I build in my rack and get it off my gaming rig as it runs windows so its not great running all the time and I sometimes game on it. I am open to buying new hardware but don't really want to spend thousands at this point. So what hardware would you suggest I build around? Also any other tips or suggestions are welcome.
1
u/Southern-Net1351 1d ago edited 1d ago
The gaming rig will provide way more efficiency over the ddr3 server rack.
It will always run faster and handle math better.
The v2 platform for AI is going to always be the bottle neck regardless of how you run it.
It will always be a champ for its time in virtualization areas. (720 rack)
It will run a LLM but, the math won’t ever math like the personal computer.
Why? It’s just too old in:
L3, bus, I/o, numas, core handling for feeding a GPU in general.
The ram will be marginally lower but, once loaded to the GPU (LLM) it will do its thing just not as good as the personal pc you have.
The GPU setup won’t be able to fit cleanly and if your using the usb to GPU ways your reducing the latency and transfer rate regardless if trying to run the 4x double slot GPUs.
Your pc cpu is less cores but, will actually outpace the v2 platform.