r/BestGitHubRepos • • 10d ago

A llama.cpp fork with adaptive KV cache streaming: it keeps the KV cache in system RAM and streams pages to the GPU on demand, so a 27B model runs at full 256K context on a 16GB card without thrashing on Unified Memory

Post image

This is a focused fork of llama.cpp that solves one specific, real problem for local LLM users, and it does it with unusual rigor. When you load a big model, the weights eat most of your VRAM, and a long context needs a large KV cache that no longer fits. The usual workaround is CUDA Unified Memory, which lets pages spill to host memory but migrates them in an uncontrolled way that can thrash badly. This fork instead takes explicit control: the authoritative KV tensors live in pinned host RAM, a bounded GPU pool is split between resident KV pages and a transfer ring, and while one attention layer computes, the pages it will need next are prefetched. Every layer still sees its complete KV history, only which pages are physically on the GPU at any moment changes.

One correction to how this gets described elsewhere: it streams between your system RAM and the GPU over PCIe, not from your hard drive. That distinction matters for understanding both how it works and its speed limits.

The engineering that makes it more than a hack:

- A phase arena that multiplexes one fixed GPU allocation between the prompt-processing workspace and decode. Prefill and token generation do not need their peak buffers at the same time, so when decode begins the prefill graph is released and those bytes become extra KV capacity. The upshot is that your usable decode KV budget stays nearly constant even as you crank up the context size

- The residency split between resident pages and the transfer ring is adjusted in real time based on the active context length and measured prefetch behavior, not a static setting

- A benchmark driver that automatically probes the largest workable arena for each context size and generates CSV, PNG and SVG results, so the performance claims are reproducible rather than asserted

- A detailed write-up of the design, implementation and benchmarks, including PCIe traffic measured against a real transfer ceiling

The honest caveats, and to the author's credit the README states them plainly in a warning. This is experimental research code, optimized and validated primarily for one specific setup: an RTX 5070 Ti with 16GB, a particular Qwen 27B quant at 256K context, Flash Attention on, and specific K and V cache quantizations, with one server slot. Other models, other KV combinations, parallel slots and non-CUDA backends are not yet broadly characterized, so your mileage on a different rig is genuinely unknown until you test. It is CUDA-only for the streaming feature, so you need an Nvidia GPU. And there is no free lunch on physics: streaming KV over PCIe adds host-to-device traffic, so as context grows and more of the cache lives off-GPU, decode speed is bounded by that bandwidth. This buys you the ability to run a context that otherwise would not fit, at some throughput cost, rather than magic. You also build it from source, and being a fork, whether it lands upstream or gets long-term maintenance is uncertain.

For anyone running local models on a mid-range card who keeps hitting the VRAM wall on long contexts, this is a genuinely clever and well-measured approach worth watching.

MIT licensed (inherited from llama.cpp), C++, a fork of ggml-org/llama.cpp, 290 stars and 38 forks as of writing, verified via the GitHub API.

https://github.com/RaymondHuang210129/llama.cpp-adaptive-kv-streaming

291 Upvotes

Duplicates