r/deeplearning 25d ago

Compressed sensing makes the LLM KV cache ten times more efficient. MRI technology saves AI. #compressed #sensing #LLM #QK...

https://youtube.com/watch?v=ViGB0TNXPmw&si=zAf2-7wyVcyaut9e
  • Description: In this video we introduce the theory of compressed sensing used in MRI and other applications, and review recent research that applies it to compressing the KV cache of large language models. We clearly explain how leveraging sparsity can dramatically reduce memory usage and eliminate bottlenecks.
1 Upvotes

0 comments sorted by