Where the arithmetic points

Put the figures on one denominator. A model with 70 billion parameters running a one-million-token context window generates roughly 320 gigabytes of key-value cache for a single user. That is four times the entire high-bandwidth memory capacity of an NVIDIA H100. The data has to go somewhere, and the nearest candidate is an NVMe drive. At PCIe Gen 5 speeds a typical enterprise drive supplies about 14 GB/s; at the sixth generation the ceiling for a standard four-lane drive rises to 28 GB/s.[1]

The CM10, which Kioxia announced on 30 July, sits in that tier: 332-layer BiCS FLASH generation ten TLC memory, and direct cold-plate liquid cooling alongside conventional air cooling. The E3.S and 9.5 millimetre E1.S form factors target dense accelerator racks where cold plates rather than fans remove the heat. It is worth naming where this sits on the ladder: the drives are sampling to select customers, price and availability have not been disclosed, and the first public demonstration is on Tuesday.[1]

The denominator behind the numbers

The performance figures are Kioxia's own measurements against its own previous generation: roughly 92 percent higher sequential read and roughly 85 percent higher random read than the CM9, with peak sequential read at 28.4 GB/s and random read at 6.29 million IOPS. The interface generation is not what separates the field; Micron's 9650 reached mass production in February 2026 and Samsung's PM1763 on 8 July. With all three vendors on the sixth generation, the difference will be found in memory density and thermal handling.[1]

Bandwidth also has a cost. The sixth generation moves from binary signalling to PAM-4 four-level encoding to reach 64 GT/s per lane and adds forward error correction for the first time in PCIe's history, which means roughly 4 to 8 nanoseconds of latency per direction. Loading model weights or staging training data, that is negligible. In inference pipelines where time to first token is measured, it is an expense to be counted separately. What would show that this tier genuinely helps is an independent measurement on a common workload.[1]