What unidentifiability means here
Hidden states are added to these models for a good reason: they absorb unobserved factors that would otherwise be attributed to the trait under study, and adding them became standard practice precisely because it is the cautious move. The finding is that the cautious move creates the problem. With hidden states present, every trait-independent scenario is congruent with an infinite set of trait-dependent ones, meaning the likelihood assigns them all the same value and no quantity of data separates them.[1]
The word to hold onto is asymptotic. This is not sampling noise that a larger phylogeny would average away; the models remain indistinguishable in the limit. The precedent is Louca and Pennell's 2020 result on congruence in diversification models, which changed practice across the field and left open a question the authors take up here: was congruence a peculiarity of that setting, or a general property? Their answer is that it is universal in this model family, and that the congruence they describe is the only type that can arise.[1]
The part that is fixable and the part that is not
The paper does not stop at the diagnosis. It offers an analytical solution that works within a congruence class, the set of models that cannot be told apart, and shows that once that solution is in place, macroevolutionary inference with these models holds up. That is a narrower and more useful claim than a general warning, and it is the difference between a methods critique that changes software and one that changes only tone.[1]
What survives is a failure the field has complained about for years without an explanation: model selection across congruence classes produces false positives at a high rate. The authors' account is that this follows from model misspecification together with what they call model proliferation, the accumulation of candidate models among which one will fit well by construction. That is a mechanism, and a mechanism is what turns a known nuisance into something a reviewer can check for.[1]
What to do with a published result
The move at risk is specific and narrow. A trait-effect result obtained by comparing a trait-dependent model against a trait-independent one drawn from a different congruence class carries exactly the false-positive risk described here. This retracts no study, and reading it as a retraction would be the wrong lesson; what changes is what the comparison licenses. For most published analyses that means stating the conclusion as compatible with the data rather than as supported against the alternative.[1]
The reform that would bite is a change in the default comparison rather than a checklist item, and it has an observable signal: whether the software packages implementing these models adopt within-class comparison, and whether the next year of papers reports which congruence class its candidate models came from. The alternative reading is that the analytical solution proves impractical for the model families biologists actually fit, in which case the honest outcome is fewer trait-effect claims rather than better ones.[1]