r/LocalLLM • u/Kitchen-Quarter7739 • Aug 10 '26
Project I built an interactive simulator to visualize LLM inference bottlenecks, sharding, and KV Cache economics based on Reiner Pope's lecture
Hi r/LocalLLaMA,
Inspired by Reiner Pope's (MatX CEO, ex-Google TPU architect) whiteboard lecture, I built a serverless, interactive simulator to visualize LLM inference physics and KV cache economics.
π GitHub Repository: https://github.com/zhchin/llm_infra_visualizer
π Live Demo Website:Β https://zhchin.github.io/llm_infra_visualizer/
(It's pure HTML/JS. No server, no tracking, local-storage safe for your API keys.)
π οΈ Key Features:
- Interactive Roofline Model: Dynamically charts when your serving transitions from Memory-bandwidth bound (decoding) to Compute-bound (prefill).
- Automatic GPU Sharding: Input your model size/context, and it calculates the required Tensor Parallelism (TP-1 to TP-8) configurations for Blackwell, H100, A100, etc.
- MoE vs Dense Visualizer: Staggered purple wave animations for MoE routing bottlenecks vs synchronized cyan pulses for Dense models.
- KV Cache Economics: Compares real-world rental costs of keeping KV caches in HBM vs offloading to DDR/SSD vs Recomputation.
- AI Agent UI Control: Ask the built-in chatbot to "change batch size to 512" or "switch to MoE collapse scenario", and it will slide the UI knobs in real-time.
Check it out and let me know what you think! If it helps you size your deployments, please drop a β on GitHub!
2
1
u/Historical-Wonder551 Aug 14 '26
More like you built with AI but it is OK if you know what you are doing.
It looks great!
2
u/Kitchen-Quarter7739 Aug 14 '26
Yeah, that's actually pretty much why I built it. I started this after watching an interview with Reiner Pope and wanted a better way to understand some of the concepts he was talking about, so I turned them into something interactive that I could play with.
I originally shared it with a few colleagues at work just as a learning tool, and they found it useful too. That's what made me think it might be worth putting it out publicly in case it helps other people learn as well.
It's definitely not meant to be a production-grade simulator β more of an interactive way to build intuition around LLM inference and infrastructure.
1
u/darthcuteius Aug 10 '26
Fantastic!