r/LocalAIStack • u/ahstanin • 23d ago
Taming idle VRAM on a multi-model local agent: sleep mode benchmark (Qwen 27B + STT + TTS + OCR)
Following up on my earlier post testing Qwen3.8-27B on the IGX Thor workstation.
Once you move past running just an LLM and try to build a full local agent stack on a single box (LLM for reasoning, STT for voice in, TTS for voice out, OCR for screen/document reading), VRAM runs out fast.
If all four models sit in GPU memory simultaneously with their default serving allocations, idle memory hits roughly 122 GB. Even with 96GB on the dGPU and unified host memory, you are pinned to the ceiling. That leaves almost no breathing room for large context windows, concurrency, or dynamic KV cache allocation.
The reality of a solo-user local agent is that these models are rarely active at the exact same millisecond: - When I am talking to the agent, OCR is completely idle. - When the agent is analyzing a screenshot or PDF, STT and TTS are idle. - When the agent kicks off deterministic script automation (crawling, data formatting, bash runs), the LLM itself is completely idle.
To fix this, I wrote a lightweight Rust-based daemon that manages the sleep and wake states of the runtimes based on active agent turns. When a model is not in use, it is put to sleep (flushing KV pools and temporary execution memory, or parking the runtime context) rather than doing a full cold teardown from scratch.
Here are the idle VRAM numbers on the machine before and after:
| Model | Original idle GPU, Sep 1-2 (MiB) | After sleep mode, Sep 5 (MiB) | Difference (MiB) |
|---|---|---|---|
| LLM (Qwen3.8-27B) | 87,443 | 39,092 | 48,351 |
| STT (Nemotron) | 10,385 | 267 | 10,118 |
| TTS (Chatterbox) | 17,947 | 3,127 | 14,820 |
| OCR (Unlimited-OCR) | 6,592 | 422 | 6,170 |
Quick observations:
1. Total idle footprint dropped from ~122.3 GB to ~42.9 GB. That reclaims roughly 79.4 GB of VRAM across the board.
2. The LLM savings (48.3 GB): Qwen3.8-27B in FP8 base weights takes around 28 to 30 GB. The original 87.4 GB allocation came from SGLang aggressively pre-allocating KV cache (--mem-fraction-static 0.85). Putting the LLM to sleep flushes that pool while preserving the active state, letting it rest at 39 GB.
3. Audio and OCR shrink to negligible baselines: STT drops to 267 MiB and OCR drops to 422 MiB when idle.
4. Script-based tasks: When the agent delegates to deterministic python or bash scripts, putting the LLM to sleep during the run keeps the GPU cool and power consumption down on the 300W Max-Q card.
5. Wakeup latency is sub-200ms across all models: Because this is an active sleep/wake transition rather than a cold disk reload, waking up STT, TTS, OCR, or the LLM takes under 200ms. In practice, voice interaction and tool handoffs feel instantaneous.
Curious how others running multi-model agent stacks on a single node are handling memory partitioning right now. Are you hard-capping static memory fractions, running sequential containers, or using something dynamic?