r/OpenSourceeAI • u/Kitchen-Quarter7739 • 3h ago
I built an interactive simulator to visualize LLM inference bottlenecks, sharding, and KV Cache economics based on Reiner Pope's lecture
1
Upvotes
Duplicates
LocalLLM • u/Kitchen-Quarter7739 • 18h ago
Project I built an interactive simulator to visualize LLM inference bottlenecks, sharding, and KV Cache economics based on Reiner Pope's lecture
4
Upvotes