What was measured against what
The abstract of the paper published in Nature states that WeatherNext Cyclones was evaluated on tropical cyclones from 2023-2025 and that its track, intensity and wind radii predictions offer an average of a day or more of lead time advantage over leading operational models. The blog post names the comparators: ECMWF-ENS for track and HWRF for intensity. Nature is providing the manuscript as an unedited version at this stage.[1]
The choice of comparator is the strongest part of this work. The benchmark is the set of operational models the centres actually run, rather than the model's own earlier version, which is why the result sits on a meaningful baseline. What is measured is forecast skill on past storms, a process metric. The path from there to fewer deaths runs through issuing warnings, evacuation and response, and this evaluation does not measure those steps.[1]
The resolution question
The model's input resolution is cells of 28 by 28 kilometres, 100 times coarser than traditional models, and the smaller WeatherNext 2-mini runs on cells of 111 by 111 kilometres. The paper's abstract says this suggests high resolution is not a strict prerequisite for state-of-the-art intensity forecasting and that coarser atmospheric data contains more intensity signal than previously recognised. The blog writes that it remains an open research question how the models produce such accurate predictions at this resolution.[1]
Two claims could explain the same observation. The first is that the coarse fields genuinely carry the intensity signal. The second is that part of the skill comes from the model's second training source, the IBTrACS database of nearly 5,000 historical storms, in which case the resulting capability sits closer to a learned climatological distribution. The two are not mutually exclusive, and the result reported in the paper does not separate them. Releasing the weights and the code makes the separation possible.[1]
The experiment that separates them
The design needed is clear: retrain the model with the cyclone-database component removed, or evaluate it on storms from basins never seen in training, and report whether the lead-time advantage survives. Because the code and weights are open, a group outside Google can do this. I expect such an evaluation to be published by August 31, 2027; if the advantage holds on held-out basins, the claim that coarse fields carry the intensity signal is strengthened, and if it falls materially, part of the skill is coming from the storm database.[1]