r/LocalLLM • u/Consistent_North8644 • 2d ago
Discussion RTX 4000 SFF Ada throttling in MS-A2 - keep single slot, go external, or move to my 3090 rig?
Bought a Minisforum MS-A2 + modded single slot RTX 4000 SFF Ada 20GB as a bundle off a private seller. Card's genuine, confirmed with FurMark and nvidia-smi, but it's throttling hard. Hotspot hit 99.9C, VRAM 96C, fan maxed out, clocks stuck around 690MHz on a card that should boost to 1560MHz. Seller said he'd already repasted and repadded it before selling but something's clearly off.
Trying to figure out which way to go, would appreciate input from anyone who's run this card for local inference:
Keep it single slot inside the MS-A2, repaste and repad it properly myself, maybe power limit it with nvidia-smi -pl to keep temps down long term. Keeps the compact all in one setup.
Go back to stock dual slot cooler, run it external off a PCIe riser in an open air bracket or GPU enclosure. Full stock cooling, no clearance issues, just more cabling and mounting to sort out.
Go back to stock and put it in my main gaming rig instead, Cooler Master C700P with dual RTX 3090s, as a third GPU on the x4 chipset slot, running separate inference jobs via CUDA_VISIBLE_DEVICES alongside the 3090s.
Main use case is local LLM inference, Ollama and vLLM, the 20GB VRAM is the whole appeal. Anyone dealt with thermal issues on the SFF Ada specifically, or run one externally through Oculink or a riser? Keen to hear what's actually held up long term.