Eigen RadarAI
Analysis

Cheaper token prices, heavier cooling loads

A Gartner model expects agentic inference spending to rise with use; TrendForce sees liquid cooling spreading, while NVIDIA unveiled a smaller model format. Efficiency gains are shifting AI infrastructure costs across software, compute and cooling.

Artificial Intelligence··Midday
An anonymous technician checks AI server racks beside transparent liquid-cooling manifolds and blue coolant lines in a bright data-centre aisle.

Token prices fall while agent bills climb

A Gartner model reported by Computerworld on 17 August forecasts that token costs will fall 95 per cent by 2030 while inference spending on agentic workflows will rise more than 5 times through 2028. Analysts Will Sommer and Sabine Zimmerhansl call the gap a token-deflation illusion in which falling unit prices hide rising consumption. Gartner's figures are forecasts from its own tokenomics model, not observed market billing; the firm puts advanced agents at up to 150 times the cost per task of a basic chatbot, with training a mid-sized agentic model 2.5 times costlier and agent inference 5 times greater. Its model prices a basic workflow near 0.05 dollars per inference token, summarisation near 0.10 dollars, complex workflows near 0.30 dollars and planning work near 0.40 dollars.[1]

Liquid cooling passes half of chip deployments

TrendForce said on 17 August that liquid cooling will cover 53 per cent of AI chip deployments in 2026, up from about 33 per cent in 2025, and approach 60 per cent in 2027. The research firm attributes the shift to individual chips passing 1 kilowatt of thermal design power; the percentages are TrendForce's own estimates rather than reported shipments. Rack-scale systems now draw hundreds of kilowatts, and the firm forecasts that NVIDIA's rack-scale shipments will double in 2026, with its Vera Rubin platform built as a fanless, all-liquid design that extends cooling to network cards, optical modules and power systems. AMD will launch its Helios rack in the second half of the year. Google already runs liquid cooling in more than 80 per cent of its AI servers.[2]

NVIDIA shrinks the model file with distillation

NVIDIA said it built the NVFP4 version of its Nemotron 3.5 Lightning model with quantization-aware distillation; according to the company the file drops from 65.85 GB to 21.19 GB and throughput can rise up to 4 times. The method has two stages: post-training quantization first produces a low-precision student from the full-precision teacher, then the student is trained against the frozen teacher with a KL divergence loss. Tables published by the company show that on one intermediate checkpoint post-training quantization gives 96.33 per cent median recovery while distillation raises it to 99.72 per cent; on the shipped checkpoint a more conservative quantization is used, and median recovery comes out at 99.24 per cent for post-training quantization and 98.97 per cent for distillation. The gains from distillation collect in agentic and coding benchmarks such as Terminal-Bench v2.1 and SWE-Bench Multilingual. All the figures are NVIDIA's own measurements, and no independent validation has been published.[3]

References

  1. News sourceComputerworldToken prices fall while agent inference bills head fivefold higher by 2028↩
  2. News sourceTrendForceLiquid cooling reaches 53 per cent of AI chip deployments this year↩
  3. News sourceNVIDIA Technical BlogNVIDIA adds distillation to 4-bit quantization↩