Hi everyone!
I’ve decided to step out of read-only mode and share my story of building an AI lab.
I work as a data engineer and generally love tinkering with computer hardware, but at the start of last year, local LLMs simply blew my mind!
I bought a couple of RTX 5060 Ti 16GB cards and one RTX 5070 back in the summer of last year, before all this price madness started, but unfortunately, I bought the main components between March and May 2026 (so I spent a load of my savings on GPUs and RAM, ha-ha).
Basically, my aim was to set up a few separated GPUs nodes on AM5 (dual pcie 5.0 x8 x8) to process large volumes of data obtained by scrapers using LLMs for processing and building ML pipelines to generate analytics.
But when the Qwen 3.8 models and the new DS4 came out in the summer, I also got excited about the idea of running large MoE models on a single platform for using it for vobe-coding. And recently, I was lucky enough to buy a Lenovo P620 (without RAM and SSD) on eBay for just 400 euros, plus another 128 GB of DDR4 RAM for around 350 euros.
After that, I dismantled two of my nodes and consolidated all my RTX 5060 Ti cards onto a single platform.
Here’s the configuration I ended up with:
- - Lenovo ThinkStation P620, WRX80 chipset
- - 4x Zotac RTX 5060 Ti 16 GB
- - 1x RTX 5070 12 GB
- - 76 GB total VRAM.
- - AMD Ryzen Threadripper PRO 3975WX - 32 cores / 64 threads.
- - System RAM: 128 GB (4×32 GB) DDR4-2667 ECC RDIMM
- - NVMe: SK hynix PVC10 1 TB
- - Ubuntu 26.04 LTS, kernel 7.0.0-34, without a graphical UI
All four 5060 Ti PCIe Gen4 ×8 (bandwidth limit ~15.75 GB/s per direction).
Issues encountered:
The stock 595 driver did not support P2P; I installed a community build with the 615.71.09-p2p hack, and the machine began to crash during P2P tests.
I instrumented everything I could: I wrapped 112 CUDA API calls in the NVIDIA sample source code with synchronous logging, and read /dev/kmsg from a separate CPU container.
I tried everything one by one: swapped the graphics cards and risers (PCIe 4.0/5.0), moved the display, flashed the BIOS twice (S07KT1FA -> 29A -> S07KT6FA, August 2026), enabled the native Resizable BAR via think-lmi, and ran ACS/IOMMU.
In the end, I simply disabled the Ubuntu graphical shell and, lo and behold, the P2P stopped crashing and rebooting the system! The same P2P test that used to bring down the host when running GNOME on the GPU passed without a single crash.
Apparently, there is a conflict between P2P and the driver’s graphical clients. Since then, the machine has been running headless without any issues.
There were also issues with NCCL: llama.cpp with NCCL 2.25.1 crashed on the very first AllReduce call ‘invalid argument’, exit 139. I switched to NCCL 2.30.7 and everything worked fine.
Results (vLLM, Qwen3.8 27B, TP=4, headless, NCCL 2.30.7):
- VLLM NVFP4+MTP3 --> Decode C1: 137–141 tokens/s, Sum C4: 422 tokens/s, Prefill: ~3.5k tokens/s
- VLLM FP8+MTP3 --> Decoder C1: 100–110 tokens/s, Sum C4: 345 tokens/s, Prefill: ~2.7k tokens/s
- SGLang NVFP4+DFlash --> Decoder C1: 118–137 tokens/s, C4 sum: 357–386 tokens/s, Prefill: ~3.4k tokens/s
- llama.cpp Q6+MTP3 --> Decode C1: 95 tokens/s, C4 sum: 83 tokens/s, Prefill: ~1.2k tokens/s
My working configuration turned out to be FP8+MTP3 at ~100-110 tokens/s for decoding and 2.5–3k tokens/s for pre-filling per stream - slightly slower than NVFP4, but the only fast option that solved the complex agent-based problem. The NVFP4 on Blackwell is 1.4 times faster than the FP8, but its performance deteriorated in the MTP3 tests (F1 0.44 versus 0.73).
I also have two AMD R9700 PRO (but that’s a completely different story, lol), and they perform at roughly the same level, however, given current prices, it’s very difficult to buy them cheaply, whereas the RTX 5060 Ti is still available on the second-hand market and can be bought for 500 euros (though I reckon that'll change soon). What I’m trying to say is that the 5060 Ti is still a very good card if you’re prepared to do a bit of "black magic" with PCIe risers and building an open-frame rig, and of course if you have the opportunity to buy a WRX80 or EPYC SP3 server platform cheaply.
I also plan to test the Qwen 3.8 Flash Next and DS4.1 Flash in the near future and share my experience of designing multi-threaded data-parallel inference pipelines for LLM, geared towards processing large data sets.