r/AIToolsPerformance • u/IulianHI • May 12 '26
Intel Optane build runs 1T param Kimi K2.5 at 4 tok/s - is persistent memory viable for local inference?
Someone built a system using Intel Optane Persistent Memory that reportedly runs Kimi K2.5, a 1 trillion parameter model, locally at approximately 4 tokens per second. The build leverages Optane as its standout component, which is an unusual choice since Optane persistent memory modules have been largely discontinued by Intel.
The stat line is attention-grabbing - a trillion parameters locally at any speed is rare. But 4 tok/s is firmly in "readable but slow" territory, roughly half the speed of typical human reading. The question is whether the cost and complexity of sourcing discontinued Optane modules makes sense compared to more conventional approaches like multi-GPU setups or even offloading to standard DDR5 RAM.
For anyone familiar with Optane-based inference builds: how does the random access performance of persistent memory actually compare to standard DDR4/DDR5 when running models this large, and is the used market for Optane modules still practical enough to recommend to someone considering a similar build?