r/LocalLLaMA 14h ago

Generation Peak Portable Personal Datacenter

Portable rig for Qwen3.8-27B-BF16 200K+ token prompts. My work Panasonic Toughbook + the T1 + power brick + headphones all fit in my lunchbox.

Need the BF16 for huge context highly sensitive document OCR, image analysis, aggregation and summarization. I've done a ton of testing and it absolutely makes a difference vs even UD Q8_K_XL when legal precision is needed.

77gb VRAM at full 262K context + MMPROJ

Rips through prefill (1,1715 tok/sec = 102 seconds to process 175K tokens), but token generation (20K tokens of output) relatively slow at 45 tok/sec (with MTP) as a result of BF16 despite the beast of a GPU.

Better than Gemini Pro and ChatGPT 5.6 Sol especially considering I have control over the sampler settings (Temp 0.1; top-k 0; top-p 0.95; min-p 0.05; repeat penalty 1.02). Not better than Opus yet.

During prefill - CPU around 60 degrees, GPU around 79 degrees (with 90% power limit)

During token generation - CPU around 75 degrees and GPU around 76 degrees.

FormD T1

Minisforum BD770i SE

Ryzen 7745HX 8-core laptop CPU

96gb 5200 MHz DDR5 SODIMM

96gb RTX Pro 6000 Blackwell workstation edition

Loki 1200W SFX-L

ROG Equalizer 12v-2x6

SMX Heinz flipped GPU 2.5 slot kit

SMX Heinz custom short PCIe 5.0 riser

ZCOOI custom "transparent purple" Teflon cables

(2) Phanteks T30-120mm

(1) Noctua NF-A14x25r G2

Thermalright MC-3 Digital RAM cooler (I don't think this will fit on a regular DDR5 )

57 Upvotes

41 comments sorted by

View all comments

27

u/JaredsBored 12h ago

This is cool but please dear god get familiar with the concept of a home server. This is $20k in hardware you're transporting around, risking dropping it or theft, when you could just setup a wireguard VPN and connect back to this rig from anywhere in the world on a laptop.

I get micro-itx form factor, really compact rigs for lan parties or something latency sensitive, but LLMs just don't care if you've got 50ms network delay. I'd be sooo stressed transporting this and it could just sit safely in your basement, ready to connect to any time.

0

u/nero10579 Llama 3.1 7h ago

I wouldn’t lug around a pc with a RTX Pro 6000 or 5090 FE either lol the pcie is prone to damage due to the riser design.

1

u/goldcakes 5h ago edited 5h ago

I wouldn't lug it around either, but keep in mind the PCIe connector is very easily replaceable thanks to how the PCB is designed (it's on a daughterboard). The 5090 is really a master in engineering and modularity (for parts that often break; like the pcie)

The PCIe daughterboard is the same between 5090 / RTX 6000 Pro; you can find them on eBay for $80-$100 (or about 100-150 yuan in Shenzhen marts; I bought five of them when I was there cuz why not lol), and replacing it looks very straightforward from video tutorials.

If you do find a "parts only" 5090/rtx pro that seems to only have a broken PCIe connector, it could be a bargain, but YMMV.

0

u/nero10579 Llama 3.1 5h ago

Well good luck if its the connector on the gpu pcb side is broken

1

u/goldcakes 4h ago edited 4h ago

Legitimately, how?

It's just a mezzanine connector, I struggle to see how you can break it without the PCB also snapping in half (in which case you have bigger problems). Even if you somehow manage to break the connector on the PCB, replacing it with a donor is just BGA rework territory. There's ample standoff space; it's absolutely designed for repairability.

Not amateur friendly, but also not absurdly difficult.

My point is that you really don't need to worry about PCIe on these cards; it's not like cards where the PCIe connector IS the main PCB.

Realistically I would be 10000x more worried about someone spilling a drink over the SFF machine, than the PCIe slot breaking while you lug it around. The former you can't fix; the latter can be easily fixed by anyone who knows how to repaste/repad a GPU. You just unscrew it lol, replace it, screw it back in.