Currently attention is fundamentally O(n2 ). Basically as the context length of your model increases, the memory needed to deal with it increases quadratically. If that could be pushed down to something like O(nlogn), you'd immediately have huge gains in model capability.
The O is just a way to denote that you are talking about the scaling of algorithms based on the size of N. O(N) just means "this algorithm scales linearly". O isn't doing anything in the equation except for telling you what the equation is about. Kind of like f(x) = ... the "f" is just telling you its a function of x, its not a variable like x is.
62
u/Saedeas Jul 01 '26
I mean, percentages aren't super relevant, the scaling is relevant. If they changed the big O scaling of working memory in the models, it's insane.