Putting the numbers on one denominator

Ai2's OlmoEarth Platform post measures a single job: wildfire-risk mapping for North America. Roughly 19,600 CPUs and 994 GPUs ran in parallel, peak network throughput passed 168 GB/s, an estimated 4,737 hours of processing came down to about 30.5 hours, and that is reported as a 155-fold speedup. Cost is given as fractions of a penny per square kilometre.[1]

Divide 4,737 hours by 30.5 and you get 155, so the speedup is the parallelism itself, and nearly all of it was realised. That makes the interesting number not 155 but 168 GB/s. A job reading dozens of terabytes of imagery is bounded by how fast bytes reach the accelerators, not by how many accelerators there are. One caution belongs here: 168 GB/s is a peak, not a sustained average. If sustained reads are far lower, the near-linear scaling may come from the CPU intensity of the postprocessing stage rather than from input and output.[1]

Which rung this sits on

This job is not on the promise rung: the post describes a completed continent-scale inference run measured in wall-clock time. But the fields that would make its cost comparable are absent. The hardware generation, the price of the CPU and GPU hours, the utilisation sustained across those 30.5 hours and the licensing terms are not stated. Without them, a figure of fractions of a penny per square kilometre stays a unit cost defined inside the vendor's own boundary, and cannot be set beside another run.[1]

The constraint ladder shows itself in the three-stage pipeline: CPU-based data acquisition, GPU inference, then CPU postprocessing again. A pipeline that places a CPU stage on either side of one GPU stage moves the bottleneck off the accelerator — and the CPU count being roughly twenty times the GPU count is precisely the measure of that. In this job the scarce resource is the ordinary compute that feeds the accelerator and gathers its output, rather than the accelerator itself.[1]