Where the two counts part
In 2023 Michael Park, Erin Leahey and Russell Funk applied the CD index to almost 45 million papers and 3.9 million patents and reported that both had grown less disruptive over the decades. The index reads a work's position in the citation network: it asks whether later work cites the paper on its own, or keeps citing what the paper itself had cited. A work with no backward citations at all takes the highest value the index can give, which is 1.[1]
Vincent Holst and colleagues at Vrije Universiteit Brussel take that last property seriously. Entries with no references are scored at the ceiling; their share of the databases falls across the period; and when the group opened the source documents behind those entries, most of them turned out to carry references after all. On that reading, part of the measured decline follows a change in the metadata more than a change in science. The group also worked through the robustness checks in the original paper and reports that none of them is built to catch this.[1]
Park, Leahey and Funk reply in the same issue. Working with the dataset, the metric and the method their critics advocate, they still find declines. And they point at the critics' own regression model, the one built to take works with no references out of the picture, which yields large declines for papers and patents at P below 0.01 in the critics' own supplementary tables.[2]
Which documents entered the sample
The reply then audits the critics' sample and finds three times as many works without references as the original data hold. It traces the excess and names it: at least 2.8 million editorials, obituaries and comments; 1.5 million books and proceedings; 254,000 product and artistic reviews. 20 per cent of that sample is not research. A keyword pass turned up 456 For Dummies guides, 50 Dr. Seuss and Curious George books, and the Captain Underpants series, none of them citing anything.[2]
Then comes the number that decides the trend: in that sample, non-research content falls from 40 per cent of the total in 1945 to 8 per cent in 2010. A drop of that size is enough by itself to produce the decline in reference-free entries that Holst and colleagues attribute to faulty metadata. Which leaves the disputed quantity moving with the document-type filter each team applied. One alternative reading deserves a hearing: part of that fall from 40 per cent to 8 per cent could reflect how far books, editorials and reviews were indexed across those 65 years, in which case it is a coverage change working through the same channel.[2]
The reply grounds its exclusions in standard practice, and that is a real norm rather than an improvisation. It is still a rule argued for while both trends are already on the table. Compare a case where the order ran the other way: 27 earthquake models were locked in 2007 and scored 11 years later, so what counted as a pass was fixed before anyone saw the outcome. Here the definition of the sample is being settled afterwards, and that is the position in which a defensible choice and a convenient one are hardest to tell apart.[2], [3]
What would settle it?
One asymmetry is visible from outside. Holst and colleagues posted their reanalysis data and code publicly on GitHub, on top of the data the original authors had already released on Zenodo. The reply says additional scripts will be made available on request. For a disagreement whose entire content is which documents entered a sample, being able to request the classification is not the same as being able to read it: the granular document-type assignment that turns 40 per cent into 8 per cent is the object under argument.[1], [2]
So there is a test, and it is cheap. If both teams apply one granular document-type classification to a single shared dataset and publish the filter code with it, the disruption trend before and after that exclusion becomes one reported comparison instead of two claims. That comparison, or a refusal to run it, should be visible by the end of 2027; the signal to watch is a published reanalysis that reports the trend under a stated exclusion rule with the code for that rule attached.[1], [2]