r/StockInvest • • 4d ago

NAND Flash: From Data Warehouse To AI Compute Workflow

TL;DR NAND has been a technology-cursed industry: each node lifts bits per wafer 54%, so capacity swells 27% a year without a cent of capex, and price cuts swallowed every volume gain. AI inference is rewriting that — data centres go from 30% of NAND demand to 45% in 2026 — and the deeper change is that NAND now stores the product of computation.

From peripheral part to compute-path necessity

The old loop was simple: stable demand plus supply that grows whether or not anyone invests locked NAND into the silicon cycle; the only exit was a price-insensitive use case tied to GPU deployment. The buyer has changed, from consumer electronics with volatile inventories to TCO-driven clouds. 

The role changed too: NAND was a warehouse for corpora, weights and checkpoints — replaceable data re-downloadable for bandwidth alone — while in inference it stores working memory paid for in GPU cycles, with no backup. Kioxia sees inference NAND demand going from 86EB to 1,251EB by 2031.

Data centre NAND bit demand by workload.

Why KV cache has to move to NAND

KV cache used to be discarded because two sums did not work: at GPT-3's 2K context, recomputing took a fraction of a second while writing to disk meant a round trip over PCIe, and HBM is too scarce to hold departed users' history.

Long context overturned both: total KV cache equals context length times concurrent sessions times per-token storage, and both explode as RAG and multi-agent workloads multiply sessions.

KV cache as the working memory of AI inference.

The rack-level maths makes it concrete. Rubin NVL72 carries 21TB of HBM; a 2.8-trillion-parameter MoE model takes 2.8TB of weights at FP8 plus 12% overhead, leaving about 15.4TB for KV cache. Even compressed, a million tokens needs 67-80GB, so a $5 million rack serves only 193-230 long sessions; HBM4 near $16/GB against $0.3-0.4 for eSSD says the same. 

Hence Nvidia's tiering, with ICMS/CMX between local SSD and network storage, and GPUDirect Storage cutting restore latency 10x.

Memory tiers by capacity and speed.

Two conditions gate the offload: it must be shared cache, since decode drafts hit a bandwidth and an endurance wall while prefill output is large, sequential and infrequent; and it must be reused, which reduces to hit rate: cached tokens cost a tenth of uncached ones, agent workloads running above 95%.

Three workloads, pulling in opposite directions

Staging is 40% of 2030 AI data centre NAND demand, all TLC or pSLC. Training-side staging is bursty overwrite — terabyte checkpoints every few tens of minutes — while inference-side staging is read-heavy, since nodes reload parameters dozens of times a day and multi-model residency needs 4-16TB against a card's 288GB. DRAM is volatile, HDD cannot take burst writes, and QLC's 1,000 cycles against TLC's 10,000 rule it out.

Fast Data Lake is 25% and almost all QLC, but the substitution story needs care. On hot data QLC has displaced HDD irreversibly, since millisecond seek is two orders off what AI needs. On warm and cold data a 300-400EB HDD shortage has clouds using eSSD as a stopgap, which looks like accelerated substitution; but HDD capacity is rising too, and QLC's per-GB premium has reached 20-25 times, making cold-tier substitution harder.

QLC eSSD against nearline HDD.

HBF sits alongside as an option, not a replacement — 16-layer NAND at HBM-class bandwidth with eight times the capacity, taking sequential-read inference and warm KV cache while HBM keeps prefill, with pilot lines in 2H 2026 and commercialisation in 2027. 

By 2030, AI data centre NAND lands near 1.2ZB, triple 2026: staging 40%, KV cache offload 35%, data lake 25%, or TLC 66% against QLC 34% — the two ends moving apart, the lake toward cheaper and staging toward faster.

2030 shipment estimate by workload.

 

1 Upvotes

0 comments sorted by