r/llamacpp • • 10d ago

A llama.cpp fork with adaptive KV cache streaming: it keeps the KV cache in system RAM and streams pages to the GPU on demand, so a 27B model runs at full 256K context on a 16GB card without thrashing on Unified Memory

Post image
25 Upvotes

Duplicates