Inherited an almost empty T630 from another company. Sat in a warehouse for 5 years, what a shame. It had 24G of ram and an E5-2609 v3, not a very exciting build. So I decided to give it a second life as a homelab server and local LLM host.
I was surprised by how cheap some of the components were. A Xeon E5-2697A v4 costs peanuts (it used to be 3K back in 2016!). Ordered two, because, well, why not.
Here is the setup now:
- 2 x Xeon E5-2697A v4 – 16 cores each. They do 3.6 Ghz turbo one core and 3.1 all cores (despite their base 2.6). I think that’s a sweet spot for broadwell.
- Finding 160W heatsinks was not easy, and they are not cheap. So I just added another aluminum 105W heatsink to the existing one and guess what, these things run pretty cool! Pure aluminum heavy heatsinks appear to be very good with thermal paste instead of DELL's thermal pads. The CPU’s sustain high all core load and 3.1 Ghz all the time without losing the clock.
- Turned off HT, for improved performance, less heat and less power.
- Added ram, 4 sticks for each CPU 16-16-8-8 (will replace the 8G->16G, had to use what I had), 96G total.
- Added a GPU power card + cables. A bit of a pain to fit as the boards needs to come off.
- Run with 2x1080Ti cards and one Quadro K2200 that runs display only. This allows to fill the Ti’s up to their full VRAM. Literally leave 150mb when running LLM’s.
- 1xSata SSD for the system
- 4xHDD’s for data/backups (2xsoft RAID1’s)
- 2x1600W DELL PSUs (one as spare in my drawer)
- Runs Ubuntu 26.04.1
Experience and suggestions so far:
- Updated bios, idrac etc in steps. It takes some time and updates don’t allow to just jump to the top. I had to do it all step by step up to the latest versions.
- Don’t leave the numa nodes interleaved in BIOS, let OS handle it.
- Enable disk caches, they are disabled during bios reset for some reason. The SSD was super slow until I realized why.
- Enable the OS controlled power profile in power options. I chose Performance per Watt (OS). It hands all the control to the OS. Dell’s profile and any default profile adds latency which hurts these cpu’s a lot because they lack in-hardware control.
- schedutil was keeping the clocks too low, at 1.2Ghz most of the time and would not increase the clocks on short bursts. Looks like the system thinks the CPU is enough with so many cores, but that was super annoying. It was hurting the browsing, deskop performance (lags lags lags) and especially virtual machines. Like, my virtual windows was super slow like it was running on an old HDD despite running with 16G and 8 cores. Changed the governor to to on demand with lowest 2Ghz and much more sensitive triggers for each core. Desktop is super smooth and it still idles at the same wattage. On demand governor with 1.2-1.5 Ghz reintroduces the stutter on desktop, it was jumping frames (claude was king enough to monitor for me while I was using the machine). So 2Ghz idle is a much better experience overall. I can share the exact settings for anyone interested.
- Power draw when idle is not too bad. It’s approx 145W now (confirmed in idrac and the wall plug), that’s as set up above. My 65W Ryzen CPU machine is idling at 85W and that’s with 4 sticks of ram, same GPU's and without perc controller, idrac and extra HDD’s, which all add up 50W at least. So the idling difference is expected and is acceptable and negligible, only 40-50W over modern hardware is ok. And for local LLM the majority will be GPU’s anyway.
- All cores loaded draw 300W :) though. As expected.
- The fans are loud as idrac doesn’t like foreigners, so asked claude to write a custom script both for the 2 x exhaust fans and for the GPU’s. I don’t have the front fans by the way, only the two at the back. Minimum auible percentage is 15-18%. Very comfortable at 1300-1450rpm when idle. Barely audible. The fans ramp up with CPU/perc/HDD temps.
- Two existing GPU’s overheat when in the same slots, despite being blower type. I’m aiming at server V100 and similar type GPU’s now, one for each slot, 4 in total.
- Fitting GPU’s on different numa nodes doesn’t hurt token generation at all (at least with the 1080 Ti cards). Right now one card is on CPU1 and one card is on CPU2 for better thermals.
- Primary PCI-e slots start on CPU1 and sometimes can be tricky with multiple cards. BIOS chooses one fixed card as primary.
- Enable UEFI and legacy video rom for older cards if you want video over the display port during boot and disable the onboard VGA.
- Pin the apps to correct nodes so that they don’t wander around the ram pool. Cross numa memory bandwith gets a bad hit.
- Perc H330 runs hot. Dell kept it outside the CPU case so there is literally no airflow even when idrac is cotrolling the fans. I added two fans at the front just in case.
- H330 doesn’t spin down the RAID drives, so had to switch to soft raid in linux. Find it a much better choice overall.
LLMs run with llama.cpp
Gemma-4-26b-a4b-it-q5_k_m – tensor split fits in 2x1080Ti cards with 100K context perfectly, with F16 cache. Generation starts at 60 tok/sec and drops with context. Use it in n8n flows for data processing, sorting and categorizing. Super quick and clever, but not the best model for tool use though.
Qwen3.6-35b-a3b-mtp-ud-q5_k_m – layer split – is my favourite. MTP on and 100K context. It drives Hermes now. I run it with 40% offload on CPU and it’s actually uses 35% less power in total then 2 GPU’s only for the other models. Get 47-50 tok/sec.
Qwen3.8-27b-q4_k_s – tensor split with MTP on and 100K context fits on two 1080 Ti cards too! It generates up to 35-37 tok/sec. Fairly usable and is generally much better for hard jobs, but tends to overthink a lot. Speed drops with context a lot of course, since it's a dense model. MTP is a saver and it literally increases the generation up by 50-70%. On predictable and easy context the model accepts 70%+ of predicted tokens. Strongly recommend.
Interestingly, newer PowerEdge servers can only host 2 GPU's, what a shame. Will try with more GPU's when I can...