What the two main variables actually are
Both of this study's main variables are estimates rather than direct observations. Language-model involvement is inferred from word frequencies through the distributional method of Liang and colleagues: the human distribution is fitted on 6,000 NSF grant abstracts starting in 2021, before the models spread, and the model distribution comes from having GPT-3.5-turbo-0125 rewrite those same abstracts. Distinctiveness is a text measure too: the average cosine distance of an abstract's SPECTER2 embedding from every abstract the same agency funded the year before, converted to a within-year percentile. What gets measured, then, is the trace each thing leaves in the prose, standing in for how much a model was used and how original the idea is.[1]
The numbers are read off those measures. Moving from the lower quartile to the upper quartile of estimated use costs 5 percentile points of distinctiveness for NSF awards and 4 points for NIH awards. For NIH proposals the same shift raises the probability of funding by 4 percentage points, while at NSF it makes no significant difference. Funded NIH grants show 5 per cent more publications, but the relationship weakens and loses significance once the count is restricted to the top 5 per cent most-cited papers. The estimates rest on grants starting in 2023 and 2024 and see only one to two years after the award, which is a short window for highly cited work to accumulate.[1]
What the investigator fixed effects buy
This is where the design earns its keep. The regressions carry grant start year, field and investigator fixed effects, so each researcher is compared with their own other proposals. That removes the time-invariant part of ability, writing habit and topic choice. The authors also add a stylistic control: they run the distinctiveness pipeline on versions of the 2021 baseline abstracts whose surface language was altered while the scientific content was held constant, and the average distinctiveness of original and rewritten abstracts comes out nearly identical. Without that control, falling distinctiveness could equally have been the models simply making sentences resemble one another.[1]
The residual gap is named by the authors themselves. Fixed effects cannot remove time-varying choices about when and on which proposal a researcher reaches for a model. A researcher may lean on one more heavily for a proposal that is already incremental, executable or close to agency priorities; in that case low distinctiveness marks the kind of project chosen rather than a shift the model produced. A second limit runs the same way, in the data: the confidential proposals come from only two large US R1 universities. That is why the authors call the results conditional correlations within investigators and decline the phrase causal estimate.[1]
The authors' reticence answers something I asked here about a different text a week ago. I judged an argument about the format of the scientific paper to stay at the level of observation, because it had no outcome measure and no comparison group. This study has both, and still declines to move to a causal claim. The distance between the two shows that in meta-science the dividing line falls on the design carrying the argument rather than on how interesting the subject is.[1], [2]
What would move this past correlation?
The first question to ask of a proxy measure is whether its own arbitrary choices are driving the result, and this paper takes that question seriously. Logistic regressions for success and negative binomial models for publication counts return the same relationships; L2 in place of cosine distance, raw distances without percentile normalisation, and field-level and subagency replications all hold. The whole detection and analysis pipeline was rerun with Claude, Gemini and Llama as reference models, and sensitivity to tokenisation, stopword handling and punctuation was tested separately. The de-identified data needed to reproduce the main figures are posted. That is close to the full set of checks a proxy measure should attract.[1]
What would carry this further is an outside change with a date on it. The paper records one: in July 2025 NIH stated that applications substantially developed by AI would not count as the applicant's original work, while NSF stressed investigator responsibility and encouraged disclosure of use. That is a break the same detector and the same design can measure on later cohorts of proposals. On the publication side there is a clean test too: only a longer window can tell whether the missing difference in NIH's most-cited work is a lag or a matter of composition.[1]