Efficiency spreads across the stack as power demand expands
New efforts to improve model, storage and processor efficiency are arriving alongside accelerating data-center electricity demand.
Artificial Intelligence··Morning
The scale of demand
A BloombergNEF projection reported by TechCrunch says data centers could consume one-fifth of U.S. electricity by 2035, with roughly 200 gigawatts of additional capacity expected over the next decade. Nearly half of that capacity is projected to support AI training and inference. The same outlook says the United States could host 64% of the world's AI chips by power demand in 2033, while new electricity demand worldwide could reach 1,935 terawatt-hours that year.[1]
Regional figures show that this growth is not only a national aggregate. The report says 34% of electricity on the PJM Interconnection and 22% of generating capacity on Texas's ERCOT grid is dedicated to data centers. BloombergNEF's new capacity estimate is 83% above its own December forecast. Those numbers do not measure computing efficiency, but they establish the fast-rising demand baseline against which each technical efficiency claim has to be understood.[1]
Efficiency claims at the model and storage layers
Google says Gemini 3.6 Flash uses up to 17% fewer tokens than its predecessor, emphasizing efficiency, latency and reliability for developers building AI agents at scale. It also introduced the cheaper Flash-Lite tier and Flash Cyber, a security model limited to governments and trusted partners. Meanwhile, flagship Gemini 3.5 Pro has not been updated since February. Within one product family, efficiency, access and top-tier performance are therefore moving on different release schedules.[2]
WEKA addresses the bottleneck at the storage layer rather than in the model itself. It says NeuralMesh 6 software and third-generation WEKApod hardware cache previously computed tokens to raise output per GPU. The company advertises 1.1 exabytes of effective capacity and 10.2 terabytes per second in a single 56-unit rack. Results from Oracle Cloud Infrastructure and CoreWeave report large gains as well, but the news record explicitly notes that WEKA supplied those benchmarks and that they have not been independently verified.[3]
Processor design and the limits of comparison
Nvidia's Vera CPU adds a third efficiency layer aimed at the processor side of agentic workloads. The company reports 176 threads from 88 custom cores, up to 1.2 terabytes per second of memory bandwidth and as much as 1.5 terabytes of capacity. Vera links to Rubin GPUs through NVLink-C2C and targets code execution, tool use, sandboxing and data pipelines. The advertised gain of up to 1.8 times over traditional x86 infrastructure also comes from Nvidia's own disclosures.[4]
Taken together, the four records show efficiency work spreading across models, caching, storage and processors while electricity-infrastructure plans also grow. Yet token use, latency, bandwidth and throughput figures from different vendors are not directly interchangeable. BloombergNEF projects aggregate demand, whereas the vendor disclosures describe claimed gains in particular systems. The selected sources do not establish how much those local improvements will offset the increase in total electricity demand.[1], [2], [3], [4]