Count the units first

The dataset contains 491,246 repeated measurements over time from 73,321 people across 25 European-ancestry cohorts. The authors accounted for that dependence in the sampling variance and assembled 195 genome-wide association summaries. They used two social-behaviour domains, parent, teacher and self reports, and ages 2-29 as moderators in one meta-regression. The design therefore puts context into the model without confusing the volume of data with the number of people.[1]

The result was 6 loci and single-nucleotide-polymorphism heritability of 2-7 per cent. Context-matched polygenic scores explained at most 1.32 per cent of additional variance in independent samples. That small share does not erase the study's value: its contribution is refusing to treat a genetic association as fixed when age and reporter change. The score cannot, however, be used as an individual explanation of a child's social behaviour.[1]

The boundary of validation

In the Nature Human Behaviour study, the authors tested the scores in 4 independent cohorts totalling 16,305 people. Alongside European-ancestry samples, they included an African-ancestry cohort of 835 people; transfer there was limited to parent-reported low prosocial behaviour at ages 7 and 8, with at most 0.95 per cent additional variance explained. Validation covers only a limited ancestry range. Measuring that boundary explicitly is a stronger methodological choice than quietly assuming universality.[1]

The strongest inference from the Nature Human Behaviour findings is that contextual matching improves prediction and that reporter differences form part of the measurement. Consistency across European-ancestry cohorts may reflect genuine biology; shared instruments and similar cohort structures may also produce part of it. The three figures for researchers and clinicians to watch are the distinct-participant denominator, the ancestry composition of the discovery sample and the additional variance explained in independent validation. A stronger next test would prespecify the same behavioural domains and reporters in non-European discovery samples and estimate the effect sizes again.[1]