Eigen RadarScience
Analysis

Prospective tests and sequence-aware risk

Nature Communications studies score earthquake models without retrospective adjustment and place repeat length within sequence structure and ancestry when estimating genetic risk.

Science··Midday
Across sunlit dry hills, varied sealed vessels line up above offset stone layers while green and violet bead sequences enter two different mechanical gates in the foreground.

An 11-year prospective earthquake test

A study in Nature Communications evaluated 27 earthquake forecasting models operated prospectively between August 2007 and August 2018. Working through the Collaboratory for the Study of Earthquake Predictability, the models produced 51,625 daily forecasts for next-day earthquakes of moment magnitude 3.95 and above in California. The forecasts were specified in advance; evaluation used the 597 earthquakes that occurred during the period and community-vetted statistical methods. About 95 per cent of the models were consistent with observation in the number test. The share fell to two-thirds in the location test, then returned to about 95 per cent in the magnitude test. Rather than giving one score in which every model is equally successful or unsuccessful, that difference shows which component of a forecast separates the models. The authors report that selected versions of established, operational-type models maintained consistency with observation. They also publish detailed recommendations, software tools and benchmarking data for groups building operational earthquake forecasting systems worldwide. The publisher notes that the available text is an unedited version released for early access and will receive further editing before final publication.[1]

The repeat sequence beyond length

The other study in the same journal examined 66 disease-associated tandem-repeat loci across 2,530 haplotypes from 1,265 unaffected donors using long-read assemblies. Applying the initial length threshold, up to 8.5 per cent of individuals carried repeats above established disease thresholds. After the researchers removed interrupted sequences, loci with uncertain disease associations and carriers inconsistent with expected inheritance patterns, the share predicted to confer disease risk fell to about 4 per cent. The analysis covers short tandem repeats as well as variable-number tandem repeats with motifs longer than 6 base pairs. It considers repeat length together with motif composition, local ancestry, linkage disequilibrium and phylogenetic relationships. According to the authors, many above-threshold expansions contain interrupting motifs or other sequence structures that reduce pathogenicity. Alleles predicted to carry risk cluster largely at adult-onset loci with reduced penetrance. Ancestry-resolved analyses also reveal population-specific repeat architectures. The donors' lack of symptoms is part of the data context in which an above-threshold length should not by itself be treated as a clinical outcome.[2]

Locking evaluation to context

The results answer different questions: one scores models predicting earthquake number, location and magnitude; the other classifies whether particular repeat alleles are expected to carry disease risk. Their shared methodological line is preservation of essential context during evaluation. In the earthquake study, model outputs were generated before the events were seen, so 11 years of outcomes cannot be fitted back into the models; number, location and magnitude also remain separate tests. In the genetics study, an initial threshold based on repeat length is reconsidered through whether the motif is interrupted, the strength of the locus association, the inheritance pattern and local ancestry. Those layers explain the difference between up to 8.5 per cent above-threshold carriage and about 4 per cent predicted risk. Neither project assumes that adding more information must always produce a higher success rate; detailed evaluation separates models or alleles more selectively. Earthquake forecast performance and genetic disease risk still cannot be converted into the same measure. Within their own fields, the Nature Communications papers show that when a test is established and which layers of information it preserves determine what its result means.[1], [2]

References

  1. News sourceNature Communications27 earthquake models were locked in advance and scored 11 years later↩1↩2
  2. News sourceNature CommunicationsIn nearly half of the above-threshold expansions, no risk is expected↩1↩2