From reference design to volume deployment
AMD revealed the Helios design in 2025 and showed it onstage at CES in January 2026. AMD's own product page describes it not as a product for sale but as an open reference design for OEM and ODM partners to build systems around. The design specifies 72 MI455X GPUs and 31 TB of HBM4, with volume deployments expected in the second half of 2026.[1], [2]
AMD's claimed advantage over Vera Rubin is also not a measured result from the same workload. The footnote says AMD Performance Labs calculated peak theoretical values across different datatypes in June 2026 and compared them with Vera Rubin NVL72, while warning that manufacturer configurations may vary. Without fixing software, precision, latency target and full-facility power, the table does not prove real training or inference speed.[1]
Five times relative to what?
The joint AMD-Cerebras design assigns prompt and long-context processing to Helios and low-latency decode and token generation to the Wafer-Scale Engine. The 'up to five times' efficiency figure is not a measured comparison with Nvidia: the companies' footnote says July 2026 modeling compared Helios plus WSE with a WSE-only configuration on Kimi 2.6 1T at a comparable interactivity point. Initial availability through Cerebras Cloud is planned for the second half of 2026.[3]
The published modeling shows the vendor-calculated efficiency difference between Helios plus WSE and a WSE-only configuration on Kimi 2.6 1T, with initial Cerebras Cloud availability planned for the second half of 2026. Production value will turn on transfer overhead, sustained throughput and facility-level energy use when context length, numerical precision and interactivity are held constant. Those measurements will show whether the modeled peak advantage becomes a repeatable gain in customer workloads.[3]