Eigen RadarAI
Analysis

Google puts 288 GB of HBM into the TPU 8i and compresses the cache down to 3 bits

At the SEMICON Taiwan 2026 memory summit, Google Cloud's Nikhil Cherian said AI infrastructure has moved from a compute-bound regime to a memory-bound one, with high-performance memory now more than 75 per cent of an AI server's hardware bill of materials. The company's answer runs on two chips and one compression method: the inference TPU 8i with 288 GB of HBM, the training TPU 8t pooling 9,600 chips, and TurboQuant, which takes the key-value cache from 32 bits to 3 bits.

Artificial Intelligence··Evening
In a bright Taipei chip exhibition, an exhibitor shows two visitors a gleaming silicon wafer above memory packages in a glass display case.

Memory's share of the server bill has passed 75 per cent, Cherian says

Nikhil Cherian, Google Cloud's senior director of supply chain infrastructure, told the SEMICON Taiwan 2026 memory summit that large models, multimodal systems, mixture-of-experts architectures and agentic AI have carried infrastructure from a compute-bound regime into a memory-bound one. In Cherian's account, high-performance memory now makes up more than 75 per cent of an AI server's hardware bill of materials, and capacity, bandwidth, latency and power draw are the bottleneck in front of scaling. Mixture-of-experts models want wide HBM capacity to hold their weights; agentic AI lengthens reasoning flows, and the key-value cache grows with context length and multi-turn dialogue. Cherian said expensive compute chips can end up waiting on memory, with the compute capacity on hand left short of full use.[1], [2]

Training and inference split onto separate TPUs

Google is pairing hardware specialisation with software optimisation and splitting training and inference across separately designed tensor processing units. The inference TPU 8i raises HBM capacity to 288 GB and on-chip SRAM to 384 MiB, so that live conversational state and the key-value cache stay inside the chip and the delay of shuttling data to external memory shrinks. For very large-scale training, the TPU 8t links 9,600 chips into a single cluster; memory pooling takes the combined HBM to about 2 PB, paired with TPU Direct Storage. Google supplied all of those hardware figures itself, and no independent measurement has been published.[1], [2]

The cache drops from 32 bits to 3 bits with TurboQuant

On the software side, Google has built TurboQuant, a quantization method that runs without retraining the model and takes a large language model's working memory from 32 bits to 3 bits. The company reports that memory use falls by as much as six times, that attention computation runs eight times faster and that accuracy does not drop; those speed-up figures have not been measured independently either. Cherian said inference demand will overtake training: as AI spreads into enterprise products, consumer services and agentic applications, data centres have to serve inference around the clock. The company is also recycling existing server parts and, through its own interface and system technologies, trying to fit older DDR4 memory back into new AI servers. TrendForce's own projection fills in the cost side: DRAM and NAND flash together make up 47 per cent of cloud providers' capital spending in 2026 and 68 per cent in 2027, and in that same projection server DRAM contract prices rise about 270 per cent for 2026 with enterprise SSD prices about 235 per cent.[1], [2]

References

  1. News sourceTrendForceGoogle answers the memory wall with 288 GB of HBM and a 3-bit cache compression↩1↩2↩3
  2. News sourceCommercial TimesMemory tops 75 per cent of the AI server bill as the TPU 8i ships 288 GB of HBM↩1↩2↩3