Which number was measured at which boundary

The partnership splits the work in two: the first pass over a prompt goes to AMD's Helios rack-scale architecture, and the token-by-token decode stage to Cerebras' Wafer-Scale Engine. The companies report five times more tokens per second per watt for the combined system, and Cerebras chief marketing officer Julie Choi says the processor's memory bandwidth is 2,000 times that of Nvidia's GPUs.[1]

Before tokens per second per watt becomes a comparable quantity, five things have to be written down: which model, which precision, which latency target, which software stack and where the power was measured. Memory bandwidth, meanwhile, is a device property rather than system throughput; 2,000 times the bandwidth does not translate into 2,000 times the output. None of that is in the announcement, so both figures stand as vendor comparisons. The competing possibility: on a decode-heavy workload the Wafer-Scale Engine's memory may genuinely dominate, and the five-times figure may hold at that boundary.[1]

The rung: announcement or shipment

What exists today is a partnership and an installation target. Cerebras says it will put AMD Helios systems into its own data centres before the end of 2026 to feed the pre-fill layer, with joint go-to-market work in the same window. There is no named customer, no installed rack, no energised capacity and no third-party measurement. Agentic coding, real-time voice and multimodal generation are the targeted workloads, not measured ones.[1]

Eliyan's round, announced the same day at a 1 billion dollar valuation for 145 million dollars, prices the same constraint from the other side: NuLink PHYs and NuGear chiplets sold for die-to-die, chip-to-chip and rack-to-rack links. Splitting pre-fill from decode across two systems moves the bottleneck onto the link between them, which is why both items share the same rung problem. Eliyan's round also discloses no performance figure and no customer name.[1], [2]

The measurement that would test the five

The five-times figure only becomes checkable once one model and one latency target are fixed and power is read at the facility boundary. If Cerebras and AMD publish such a measurement by 31 December 2026 the number can be tested; if they do not, the one comparable quantity by that date is the count of Helios racks installed in Cerebras data centres.[1]