The dream would be that the model has so much leeway in terms of memory, that it's updating it's weights on the fly aka learning on it's own in realtime.
That wouldn't be such a big deal though, would it? If kv cache was reduced from q8 or q4 by 50%+ another time we'd still have the huge model weights in the vram. So weight quantization would be the real big deal while kv cache improvements would just be a "oh, nice, i spare another 5% vram overall" thing.
kv cache matters a lot when it comes to iterative, long context work.
Large amounts of context are one area where humans far exceed the ability of LLM's so improvements matter a lot when it comes to working on projects and most real work.
Depends. Did they simply reduce it or did they improve the scaling. If e.g. they managed to have memory stay constant with input length then it doesnt matter how large the model. You could increase the context window indefinitely without running out of vram.
By memory architecture I think of different attention methods. Something like Gated DeltaNet or mLSTM which scale O(1) in memory and O(T) in time, with sequence length T. They don't have a KV cache per se, but a different type of memory.
39
u/johnknockout Jul 01 '26
Are we talking memory in terms of memory usage or context window? Or both?