What the five questions actually asked
The manual arm is the one that bites. A curated sample of 127 quantum computing papers from the past five years was scored against five questions — code availability, environment specification, documentation, hardware description, executability — and the last of those is the only one that requires anyone to press run. Of those 127 papers, 24.4 per cent supplied code the team could even attempt, and 64.5 per cent of that code failed in a clean environment. Multiply those through and roughly nine papers in ten in that sample left nothing an outside reader could run.[1]
The second arm is much larger and much thinner. A program screened nearly 5,000 papers for the same indicators and returned a code availability rate of 26.8 per cent, close enough to the manual 24.4 per cent that the authors treat it as corroboration. It also found that about one-third of the papers with accessible code carry no machine-readable environment specification. What it did not do is attempt execution on any of them.[1]
Two rates that measure different things
So the paper carries one number about whether code exists and one number about whether code exists and runs, and the agreement between 26.8 per cent and 24.4 per cent tells us only that the automated screen found artefacts at about the manual rate. It says nothing about the 64.5 per cent execution failure, because the automated screen never tried to run anything. The honest reading is that the large sample corroborates availability and leaves executability resting entirely on 127 papers. A gentler explanation is available and should be stated: the two arms were designed to answer different questions, and the authors say so, so the gap is a scoping decision rather than a flaw hidden in the method.[1]
The denominator deserves the same care. 127 is a curated sample, not a random draw from the field, and a curated sample of papers chosen for this kind of audit can easily sit above or below the field average depending on how it was assembled. That does not weaken the finding for those 127 papers — the code either ran or it did not — but it does limit what the percentage can be carried over to. The claim the evidence supports is that in a sample of this size, executability failed most of the time it was tested.[1]
What would count as the signal beating again?
Mauerer's own account weakens the easy villain. His group ran a smaller version of this four or five years ago, found the same bad picture, and expected better numbers by now; he attributes much of the gap to hardware that varies from day to day rather than to authors cutting corners. Fred Chong at the University of Chicago pushes further and says a field this young should spend its effort on innovation, noting that in conventional computer science it took decades for researchers to start insisting on reproduction standards for each other. Both positions are compatible with the measurement; neither changes what happened when the code was run.[1]
That gives a testable expectation rather than a verdict. If a comparable manual sample of papers published after this preprint is scored with the same five questions, execution step included, and no major venue has started requiring artefact evaluation in the meantime, the share of papers with runnable code should still come in under a quarter when someone reports it by the end of August 2027. A published rerun of this framework, giving both the availability rate and the execution failure rate, is the signal to watch. Until a measurement like that beats again, the field has one audit, not a trend.[1]