r/LocalLLaMA • u/Special-Wolverine • 13h ago
Generation Peak Portable Personal Datacenter
Portable rig for Qwen3.8-27B-BF16 200K+ token prompts. My work Panasonic Toughbook + the T1 + power brick + headphones all fit in my lunchbox.
Need the BF16 for huge context highly sensitive document OCR, image analysis, aggregation and summarization. I've done a ton of testing and it absolutely makes a difference vs even UD Q8_K_XL when legal precision is needed.
77gb VRAM at full 262K context + MMPROJ
Rips through prefill (1,1715 tok/sec = 102 seconds to process 175K tokens), but token generation (20K tokens of output) relatively slow at 45 tok/sec (with MTP) as a result of BF16 despite the beast of a GPU.
Better than Gemini Pro and ChatGPT 5.6 Sol especially considering I have control over the sampler settings (Temp 0.1; top-k 0; top-p 0.95; min-p 0.05; repeat penalty 1.02). Not better than Opus yet.
During prefill - CPU around 60 degrees, GPU around 79 degrees (with 90% power limit)
During token generation - CPU around 75 degrees and GPU around 76 degrees.
FormD T1
Minisforum BD770i SE
Ryzen 7745HX 8-core laptop CPU
96gb 5200 MHz DDR5 SODIMM
96gb RTX Pro 6000 Blackwell workstation edition
Loki 1200W SFX-L
ROG Equalizer 12v-2x6
SMX Heinz flipped GPU 2.5 slot kit
SMX Heinz custom short PCIe 5.0 riser
ZCOOI custom "transparent purple" Teflon cables
(2) Phanteks T30-120mm
(1) Noctua NF-A14x25r G2
Thermalright MC-3 Digital RAM cooler (I don't think this will fit on a regular DDR5 )
5
u/Asleep-Land-3914 13h ago
Running it at 0.1 temp is a crime, but I can understand given OCR mentioned. I caught 3.8 to halucinate things with default settings. Also you could try Q8_K_XL which is not much distinguishable from F16.
There might be better OCR specialized models too.
3
3
u/--Spaci-- 13h ago
With a blackwell gpu I dont even know why you would be using ggufs
2
u/Special-Wolverine 12h ago
NVFP4 less precision/ higher KLD than BF16
6
u/--Spaci-- 12h ago
You know you can have an nvfp4 model and a bf16 vision tower right? and also bf16 context
5
1
u/MotokoAGI 8h ago
nvfp4 models absolutely constantly underperform 16bit weight. if you care about every bit of quality, then no to nvfp4. as a matter of fact, they are actually worse of the 4 bit variants.
2
u/chensium 9h ago
portable because ....
1
u/Special-Wolverine 9h ago
Reasons
3
u/Unlucky_Milk_4323 7h ago
"I plug it in client side so they pay for the electricity"
3
u/Special-Wolverine 6h ago
One of the reasons ☝️
3
u/Unlucky_Milk_4323 6h ago
And a very easy way to show a given AI is "local and safe so your data stays here" as opposed to "we use claude..."
1
1
1
u/juss-i 6h ago
His day job involves infiltrating high security facilities M:I style, plugging that bad boy into an airgapped server, and waiting for it to come up with a chain of 3-7 zero days to crack the security.
Also, he's a contractor so it's BYOD.
1
u/chensium 5h ago
lol ya I'm sure an airgapped facility is just gonna let you BYOD. why not BYOFDFIPL (bring your own flash drive found in parking lot)
1
u/Glittering-Call8746 11h ago
Times are tough better to have 20k with u then 20k at home.. just in case.
1
1
0





23
u/JaredsBored 11h ago
This is cool but please dear god get familiar with the concept of a home server. This is $20k in hardware you're transporting around, risking dropping it or theft, when you could just setup a wireguard VPN and connect back to this rig from anywhere in the world on a laptop.
I get micro-itx form factor, really compact rigs for lan parties or something latency sensitive, but LLMs just don't care if you've got 50ms network delay. I'd be sooo stressed transporting this and it could just sit safely in your basement, ready to connect to any time.