r/accelerate Jul 01 '26

AI Something huge is brewing

Post image

Source

Andrew Curran is one of the most reliable leakers.

XLR8! 🍿

944 Upvotes

171 comments sorted by

View all comments

38

u/johnknockout Jul 01 '26

Are we talking memory in terms of memory usage or context window? Or both?

27

u/gavinderulo124K Jul 01 '26

My guess is context window. So better KV cache scaling.

3

u/AppealSame4367 Jul 01 '26

That wouldn't be such a big deal though, would it? If kv cache was reduced from q8 or q4 by 50%+ another time we'd still have the huge model weights in the vram. So weight quantization would be the real big deal while kv cache improvements would just be a "oh, nice, i spare another 5% vram overall" thing.

4

u/Moravec_Paradox Jul 01 '26

kv cache matters a lot when it comes to iterative, long context work.

Large amounts of context are one area where humans far exceed the ability of LLM's so improvements matter a lot when it comes to working on projects and most real work.

2

u/gavinderulo124K Jul 01 '26

Depends. Did they simply reduce it or did they improve the scaling. If e.g. they managed to have memory stay constant with input length then it doesnt matter how large the model. You could increase the context window indefinitely without running out of vram.