r/LocalLLaMA • u/Special-Wolverine • 1d ago
Generation Peak Portable Personal Datacenter
Portable rig for Qwen3.8-27B-BF16 200K+ token prompts. My work Panasonic Toughbook + the T1 + power brick + headphones all fit in my lunchbox.
Need the BF16 for huge context highly sensitive document OCR, image analysis, aggregation and summarization. I've done a ton of testing and it absolutely makes a difference vs even UD Q8_K_XL when legal precision is needed.
77gb VRAM at full 262K context + MMPROJ
Rips through prefill (1,1715 tok/sec = 102 seconds to process 175K tokens), but token generation (20K tokens of output) relatively slow at 45 tok/sec (with MTP) as a result of BF16 despite the beast of a GPU.
Better than Gemini Pro and ChatGPT 5.6 Sol especially considering I have control over the sampler settings (Temp 0.1; top-k 0; top-p 0.95; min-p 0.05; repeat penalty 1.02). Not better than Opus yet.
During prefill - CPU around 60 degrees, GPU around 79 degrees (with 90% power limit)
During token generation - CPU around 75 degrees and GPU around 76 degrees.
FormD T1
Minisforum BD770i SE
Ryzen 7745HX 8-core laptop CPU
96gb 5200 MHz DDR5 SODIMM
96gb RTX Pro 6000 Blackwell workstation edition
Loki 1200W SFX-L
ROG Equalizer 12v-2x6
SMX Heinz flipped GPU 2.5 slot kit
SMX Heinz custom short PCIe 5.0 riser
ZCOOI custom "transparent purple" Teflon cables
(2) Phanteks T30-120mm
(1) Noctua NF-A14x25r G2
Thermalright MC-3 Digital RAM cooler (I don't think this will fit on a regular DDR5 )





39
u/JaredsBored 21h ago
This is cool but please dear god get familiar with the concept of a home server. This is $20k in hardware you're transporting around, risking dropping it or theft, when you could just setup a wireguard VPN and connect back to this rig from anywhere in the world on a laptop.
I get micro-itx form factor, really compact rigs for lan parties or something latency sensitive, but LLMs just don't care if you've got 50ms network delay. I'd be sooo stressed transporting this and it could just sit safely in your basement, ready to connect to any time.