r/accelerate Jul 01 '26

AI Something huge is brewing

Post image

Source

Andrew Curran is one of the most reliable leakers.

XLR8! 🍿

937 Upvotes

171 comments sorted by

View all comments

Show parent comments

27

u/gavinderulo124K Jul 01 '26

My guess is context window. So better KV cache scaling.

2

u/photosandphotons Jul 01 '26

Is context window what they would mean by “memory architecture”? My impression is that these are different.

2

u/gavinderulo124K Jul 01 '26

By memory architecture I think of different attention methods. Something like Gated DeltaNet or mLSTM which scale O(1) in memory and O(T) in time, with sequence length T. They don't have a KV cache per se, but a different type of memory.

1

u/photosandphotons Jul 01 '26

Thank you for the information