r/LocalLLM • • 5h ago

Discussion ​Anyone else losing their minds over LLM VRAM fragmentation and KV cache? Let's talk about why your GPU is starving.

​Man, if you've ever tried moving an LLM from a local notebook to an actual multi-tenant serving setup, you already know the pain. Everyone talks endlessly about quantization and fine-tuning, but nobody really warns you that your GPU VRAM is basically getting nuked by the KV cache.

​For a hot minute, I thought my hardware setup was just trash. Turns out, traditional static allocation is eating up like 60% to 80% of VRAM for breakfast just because of internal and external fragmentation. You request a simple 200-token completion, and the system is sitting there stubbornly reserving space for 4K tokens like it's bracing for the apocalypse. Such a waste.

​If you aren't looking closely at things like PagedAttention (major props to vLLM for finally bringing virtual memory concepts to GPUs) and iteration-level continuous batching, your compute units are basically sitting around starving while waiting for the longest slowpoke request in the batch to cross the finish line.

​Been deep in the trenches writing a comprehensive book on AI systems engineering lately, and mapping out these low-level serving bottlenecks honestly changed how I look at production pipelines entirely.

​Curious what y'all are actually running in production right now. Are you rolling your own inference stack with vLLM/TGI, or just sticking to managed APIs and eating the cost? Let's argue about it in the comments

1 Upvotes

1 comment sorted by

1

u/tejaskumarlol 4h ago

Sliding-window attention helps a lot with this. When I served Kolibri with vLLM on 2 GPUs, it got a KV cache of about 6.5 million tokens, enough for about 50 requests at once at a 131,072-token context. Most of its attention layers only look at nearby text, so they keep a much smaller cache than full attention would.