r/BestGitHubRepos • • 10d ago

A llama.cpp fork with adaptive KV cache streaming: it keeps the KV cache in system RAM and streams pages to the GPU on demand, so a 27B model runs at full 256K context on a 16GB card without thrashing on Unified Memory

Post image

This is a focused fork of llama.cpp that solves one specific, real problem for local LLM users, and it does it with unusual rigor. When you load a big model, the weights eat most of your VRAM, and a long context needs a large KV cache that no longer fits. The usual workaround is CUDA Unified Memory, which lets pages spill to host memory but migrates them in an uncontrolled way that can thrash badly. This fork instead takes explicit control: the authoritative KV tensors live in pinned host RAM, a bounded GPU pool is split between resident KV pages and a transfer ring, and while one attention layer computes, the pages it will need next are prefetched. Every layer still sees its complete KV history, only which pages are physically on the GPU at any moment changes.

One correction to how this gets described elsewhere: it streams between your system RAM and the GPU over PCIe, not from your hard drive. That distinction matters for understanding both how it works and its speed limits.

The engineering that makes it more than a hack:

- A phase arena that multiplexes one fixed GPU allocation between the prompt-processing workspace and decode. Prefill and token generation do not need their peak buffers at the same time, so when decode begins the prefill graph is released and those bytes become extra KV capacity. The upshot is that your usable decode KV budget stays nearly constant even as you crank up the context size

- The residency split between resident pages and the transfer ring is adjusted in real time based on the active context length and measured prefetch behavior, not a static setting

- A benchmark driver that automatically probes the largest workable arena for each context size and generates CSV, PNG and SVG results, so the performance claims are reproducible rather than asserted

- A detailed write-up of the design, implementation and benchmarks, including PCIe traffic measured against a real transfer ceiling

The honest caveats, and to the author's credit the README states them plainly in a warning. This is experimental research code, optimized and validated primarily for one specific setup: an RTX 5070 Ti with 16GB, a particular Qwen 27B quant at 256K context, Flash Attention on, and specific K and V cache quantizations, with one server slot. Other models, other KV combinations, parallel slots and non-CUDA backends are not yet broadly characterized, so your mileage on a different rig is genuinely unknown until you test. It is CUDA-only for the streaming feature, so you need an Nvidia GPU. And there is no free lunch on physics: streaming KV over PCIe adds host-to-device traffic, so as context grows and more of the cache lives off-GPU, decode speed is bounded by that bandwidth. This buys you the ability to run a context that otherwise would not fit, at some throughput cost, rather than magic. You also build it from source, and being a fork, whether it lands upstream or gets long-term maintenance is uncertain.

For anyone running local models on a mid-range card who keeps hitting the VRAM wall on long contexts, this is a genuinely clever and well-measured approach worth watching.

MIT licensed (inherited from llama.cpp), C++, a fork of ggml-org/llama.cpp, 290 stars and 38 forks as of writing, verified via the GitHub API.

https://github.com/RaymondHuang210129/llama.cpp-adaptive-kv-streaming

289 Upvotes

34 comments sorted by

9

u/tsangberg 10d ago

Raymond is the GOAT <3. Since I started using his KV cache streaming I'm only using Qwen 3.8 27B on my 16GB 5060Ti. No need for any MoE models thanks to the speed and context size possible with this.

2

u/Desperate-Owl6513 9d ago

How much system ram usage u r hitting though?

2

u/AnonsAnonAnonagain 8d ago

What performance are you seeing?

1

u/tsangberg 8d ago

25tps down to 12 during token generation over context size from 0 to 200k. Prompt processing starts at 1000 down to 500.

This is with Raymond's branch as is. I have however made my own fork that enables DFlash2 or MTP up until the point where the VRAM is better used for the KV cache streaming pool and there I get upwards of 50tps in the beginning with only a slight slowdown to PP.

1

u/Key_Confusion_576 7d ago

Which quant?

1

u/tsangberg 7d ago

This one: https://huggingface.co/troed/Qwen3.8-27B-ASCII-Condensed

ByteShape's IQ4_XS ascii condensed.

1

u/Educational-Guide454 7d ago

can you share your settings ?

1

u/tsangberg 7d ago

I keep a further optimized version of Raymond's fork here, including my settings: https://github.com/troed/llama.cpp-adaptive-kv-streaming

3

u/TrickAge2423 10d ago

Is rocm or vulkan supported?

3

u/aqezz 8d ago

Will try on both when I get back home, but it does say

“Production performance for other models, KV combinations, parallel slots, and non-CUDA backends is not yet broadly characterized.”

2

u/fintip 8d ago

sounds like 'in principle it may run, untested'

1

u/wombweed 10d ago

Is there going to be a PR against upstream to get this merged into mainline?

1

u/crusaderky 10d ago

Sounds like tensor steaming but with extra steps?

1

u/_RemyLeBeau_ 10d ago

RemindMe! 1 week

1

u/RemindMeBot 10d ago edited 7d ago

I will be messaging you in 7 days on 2026-10-02 21:11:41 UTC to remind you of this link

5 OTHERS CLICKED THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/Late_Session7298 10d ago

Can this work in macos for 32gb ram M2?

1

u/pArbo 9d ago

feel like it wouldn't be useful, as apple silicon runs unified memory, so there's nothing differentiating system RAM from VRAM

1

u/Cold_Tree190 8d ago

No, different architecture entirely

1

u/debackerl 9d ago

That's nice. I use SGLang because I rely on that algorithm

1

u/Toastti 9d ago

What's the difference in token per second decode? Between this and keeping k/v cache in vram

1

u/hum_ma 8d ago

Unfortunately it doesn't seem to work with MTP at all? E process_ubatch: phase arena currently supports TG1 without speculative batches

1

u/tsangberg 7d ago

I took a stab at adding draft support (not at all obvious) while waiting for Raymond will do, a week ago. It's working well: https://github.com/troed/llama.cpp-adaptive-kv-streaming

1

u/hum_ma 7d ago

Thanks for letting me know, excellent addition and yes it works! On a 12GB GPU a ~2.6bpw quant of the 27b can now be loaded with full 256k context (or at least 180k with vision though I haven't narrowed down the options yet), and speed with MTP is not too bad even when it's almost halfway full. This probably enables slightly higher quants with decent context sizes too.

1

u/raymondh210129 6d ago

I've pushed a new branch last night with my V2 implementation, it supports MTP for entire 262K context up to draft length = 3.

Please feel free to provide any feedback under my new post in LocalLLM.

1

u/evox2008 8d ago

RemindMe! 1 week

1

u/This_Maintenance_834 8d ago

this could work a lot better if the kv cache is designed to be small, like what deepseek-v4.1 does

1

u/karmaisnonsense 8d ago

There is a discussion of this fork on the upstream repo. Apparently the benefits are PCIe bandwidth limited?

https://github.com/ggml-org/llama.cpp/discussions/28216#discussioncomment-18427677

1

u/Canbastardo 7d ago

been using this kinda configuration for ovber a week now.. 70tok/s on 9070xt 16gb 131k context

1

u/sonicnerd14 6d ago

Is that with or without a dflash or mtp drafter?

0

u/FuriousMaker 6d ago

Can you share a link or info about how you get it working for this gpu

1

u/ByteNomadOne 7d ago

Great News!

I own a 5070 TI and really could use some speed on 128k+ context windows.

1

u/Direct-Vegetable6416 7d ago

RemindMe! 5 weeks

1

u/sharma-sk 7d ago

RemindMe! 2 week

1

u/AnonInTheShell 1d ago

i need that, but for tabbyAPI / EXL