What exactly was put into the simulation

Carolyne Jie Huang, Samuel Pawel, Kimberley Elaine Wever, Benjamin Victor Ineichen and Rachel Heyard built 648 scenarios by varying effect size, heterogeneity, animal sample sizes and the number of pooled animal studies. Nine metrics were tested in each: the significance criterion, meta-analysis, the replication Bayes factor, weighted and unweighted Edgington methods, the golden skeptical p-value in three versions, and the controlled skeptical p-value.[1]

The results line up as follows. Where the true effect was null and there was no heterogeneity, every metric except meta-analysis and the replication Bayes factor held its false positive rate at the theoretical level. As heterogeneity rose, all of them became liberal. Meta-analysis gave the highest power and its false positive rate climbed to 40 per cent. The controlled skeptical p-value and the weighted Edgington method stayed relatively consistent across scenarios.[1]

The bound you cannot spend your way out of

The most useful result is this: translation power is set by whichever side, animal or human, carries the weaker evidence. That means a large, well-run human study cannot rescue the situation when the animal literature is thin. The reverse holds too; a strong animal base does not carry an underpowered human study.[1]

That the degradation arrives with heterogeneity suggests the ranking depends less on the family a metric belongs to than on how variable the scenario is. Another reading exists: the ranking may be conditional on the simulation's parameters. The authors write this themselves; the parameters came from a single meta-analysis, at most five animal studies were pooled, the human sample size was fixed at 107, and the work addressed statistical rather than biological translation. In a different parameter region the order of the metrics could change.[1]

What would make this bite

The finding here is about which rule you write down before you look at the data. Its place of use follows: the choice of metric belongs in a preclinical-to-clinical plan, right next to the sample size. A metric chosen afterwards becomes as selectable as the study itself.[1]

The testable expectation: if a translational programme declares its metric and its heterogeneity assumption before the human study begins, the gap between how often results are published and how often they are confirmed should narrow. The signal to watch is a registered preclinical-to-clinical programme whose protocol carries the name of its metric.[1]