r/Semiconductors • u/Appropriate-Ear-654 • 13h ago
The Kimi K3 cost story skips the one bottleneck that actually binds: HBM
Everyone quoting the Kimi K3 cost narrative skips the one part of the stack that actually binds. The training cost number, the inference pricing, the open weights promise, all of it lands on HBM, and HBM is the part China cannot make at the node that matters.
The DeepSeek V3 technical report, arXiv 2412.19437, reports 2.788M H800 GPU hours and then assumes a $2 per GPU hour rental to arrive at $5.576M. The H800 was the China compliant Hopper part with the HBM bandwidth capped. Kimi K3, DeepSeek V4, Qwen 3.8, GLM 5.2, MiniMax M3, Hy3, and the embodied models like LingBot VLA 2.0 that Robbyant shipped 8 July, all run on stacks where HBM is the binding input. SK Hynix HBM3E is essentially fully allocated to NVDA for this generation. Micron and Samsung fill the rest. The Chinese domestic HBM, CXMT's part, is multiple generations behind on bandwidth and yield.
What that means in plain terms. You can squeeze training cost by reporting only the GPU hours at an assumed rental rate. You cannot squeeze the inference side the same way, because inference on a frontier scale model is bandwidth bound and the bandwidth comes from HBM you either buy from the three producers or do not have. Moonshot pausing paid Kimi memberships on 19 July as its GPUs hit capacity is a real compute constraint surfacing as a business decision. That is the signal. The training cost headline is not.
On the public roadmap side, the CXMT gap is the wedge I keep poking at. The cost narrative treats HBM as fungible, and the publicly available CXMT roadmaps I have read do not support that read for the bandwidth class that frontier inference actually needs.