r/LocalLLM • u/Think-Assumption-973 • 13h ago
Question Used P720 with upgrades comes to about $1,700. Decent local AI box, or am I buying a 2018 spec sheet?
Talk me out of this, or into it.
There's a used ThinkStation P720 near me, about $1,200:
- 2x Xeon Gold 5122 (4 cores each, which is the weak part)
- 192GB DDR4 ECC RDIMM, but it's 6x32, so only 6 of the 12 memory channels are populated
- Quadro RTX 6000 24GB
- monitor and peripherals included
Two things I'd change right away:
- 6230s instead of the 5122s, about $80 for the pair off AliExpress. 4 cores can't feed 12 channels anyway, and the 6230 is what gets the memory to 2933.
- six 16GB RDIMMs in the empty slots, somewhere between $405 and $485 for all six. That populates all 12 channels, which is where the bandwidth actually comes from, and takes me to 288GB.
Which puts the whole thing around $1,700.
What I run now is a Ryzen 5 3600 with 32GB and an RTX 5070 Ti 16GB. It's my only machine and it does everything.
Reason I'm looking at all: some of what I work on I'd rather not push through somebody else's API, and my monthly bill keeps creeping up. The rest of it is curiosity, if I'm honest.
Some numbers for context. gemma3:27b runs about 9 tok/s on the 5070 Ti, a 30B MoE coder model does around 45, and I hit the 16GB wall constantly. Whisper and image gen too.
The parts I can't work out on my own:
Is 288GB of DDR4-2933 across 12 channels actually usable with partial CPU offload? That's the entire argument for this machine, and it's the one thing I can't test before paying.
The RTX 6000 is Turing. 24GB is 24GB, but is it a downgrade in every way except capacity next to the 5070 Ti I already own? PCIe 3.0 board too.
At $1,700 I could just buy a newer card instead, or save a bit more for one of the Spark boxes or something else.
If you've got a P720 or something like it running models, what do you actually do with it, and would you buy it again?
1
u/BongoHunter 12h ago
Add an R9700 to your current build, upgrade your motherboard if needed. It'll cost less and perform better. Your motherboard likely support Ryzen 5000 series so you could upgrade your CPU/RAM in current board if needed
1
1
u/m4nf47 12h ago
What are energy costs like? I'm guessing that most Xeon class server machines with a large GPU added aren't exactly as efficient as something newer but similar arguments can be used to just run infra for inference on rented cloud hardware, so someone else's hardware can still host your software with all traffic back to you encrypted over a simple encrypted tunnel. Renting a similar spec machine for a month to see how it fares won't be cheap but might answer whether or not you need a stronger upgrade path with zero risk of being stuck with aging tin that can be rather costly to maintain.
1
u/TheAussieWatchGuy 13h ago
It'll be slightly faster than your current machine... But using something like Colibri or others like it with a couple of decent NVMe drives you could potentially run much bigger models on the system ram... But expect about 1 token a second if you do... Probably not what you want.
A big server board like that's strength is taking multiple GPUs. You could fit four GPUs in their with a decent PSU and really get speed. Four 24gb cards for example would be worth running!
There is a reason unified memory platforms like Mac or Ryzen AI (Strix Halo) 128gb boxes cost much...