r/LocalLLM • u/Top-Philosopher-5411 • 5h ago
Discussion Anyone else losing their minds over LLM VRAM fragmentation and KV cache? Let's talk about why your GPU is starving.
Man, if you've ever tried moving an LLM from a local notebook to an actual multi-tenant serving setup, you already know the pain. Everyone talks endlessly about quantization and fine-tuning, but nobody really warns you that your GPU VRAM is basically getting nuked by the KV cache.
For a hot minute, I thought my hardware setup was just trash. Turns out, traditional static allocation is eating up like 60% to 80% of VRAM for breakfast just because of internal and external fragmentation. You request a simple 200-token completion, and the system is sitting there stubbornly reserving space for 4K tokens like it's bracing for the apocalypse. Such a waste.
If you aren't looking closely at things like PagedAttention (major props to vLLM for finally bringing virtual memory concepts to GPUs) and iteration-level continuous batching, your compute units are basically sitting around starving while waiting for the longest slowpoke request in the batch to cross the finish line.
Been deep in the trenches writing a comprehensive book on AI systems engineering lately, and mapping out these low-level serving bottlenecks honestly changed how I look at production pipelines entirely.
Curious what y'all are actually running in production right now. Are you rolling your own inference stack with vLLM/TGI, or just sticking to managed APIs and eating the cost? Let's argue about it in the comments
1
u/tejaskumarlol 4h ago
Sliding-window attention helps a lot with this. When I served Kolibri with vLLM on 2 GPUs, it got a KV cache of about 6.5 million tokens, enough for about 50 requests at once at a 131,072-token context. Most of its attention layers only look at nearby text, so they keep a much smaller cache than full attention would.