What the scan actually contrasted
The experiment is compact and the numbers matter. In the scanner, 23 people listened to 72 primate calls and sorted them by species: 18 each from humans, chimpanzees, bonobos and macaques, taken from 6 to 8 different individuals per species. Each clip ran 750 ms, and the team deliberately left loudness unequalised so the recordings would stay natural. That last choice is honest and costly at once: naturalness is preserved, and one of the acoustic dimensions that could drive a response is left free to vary.[1]
The region under test was not drawn from these 23 brains. The temporal voice areas were localised in a separate sample of 98 participants, contrasting human vocal sounds against animal sounds, nature sounds, music and noise, at a false-discovery-rate threshold of 0.05 with clusters above 10 voxels. That order of operations is the strongest part of the design. The area was fixed before the contrast of interest was run, so the effect cannot be a by-product of hunting for wherever chimpanzee calls happened to win.[1]
Two rulers that give the same order
Then comes the part the authors flag themselves. In their words, phylogeny and acoustics were at least partly confounded in the factorial design, and they add that dedicated synthetic and natural vocal material would disentangle the issue. Four species, ranked by evolutionary distance from us, come out ranked in roughly the same order by acoustic distance from a human voice. Any contrast built on that set inherits both rankings at once.[1]
The response to that is three models, and they are more careful than most. One adds mean fundamental frequency and mean energy as covariates of no interest. One adds a single Mahalanobis distance computed across 16 acoustic parameters. One adds six features chosen by discriminant analysis: loudness, intensity, change in spectrum, F2 bandwidth contour, F0 power and the intensity contour difference. What survives adjustment is the contrast the measured acoustics do not explain. The scope of the claim follows from that: the chimpanzee effect is bounded by whatever acoustic similarity those 16 parameters and six features failed to capture. The covariates may already cover the dimensions the voice areas care about, in which case what is left tracks phylogenetic proximity more than leftover sound.[1]
The measurement that would separate them
This is where the usual reflex, running more participants, buys the least. Precision on the chimpanzee contrast rises with the sample, and the contrast stays confounded, because the overlap lives in the stimulus set and not in the sample. The authors say plainly that they cannot rule out that including more primate species would have influenced the results. The axis that needs extending is the species axis.[1]
The measurement that would settle it is quite specific. Take the same species-categorisation task and the same Mahalanobis covariate, and add to the set at least one primate whose acoustic distance from a human voice disagrees with its phylogenetic distance: a distant relative that happens to sound close, or a close one that sounds far. If the anterior superior temporal gyrus follows the acoustic ordering there, the region is doing acoustics; if it follows the phylogenetic ordering, it is doing something else. Until a set like that is scanned, what we hold is a well-controlled result with its boundary written down, which is worth more than a clean headline.[1]