The silenced number
The Hume AI team's study, published on Hugging Face, runs 11 open-source speech recognition models through three probes. The second one is very plain: the numbers in an audio file are silenced, and the model is then asked to write down what it hears. The number has gone from the audio, so a correct output should contain no number at all. On LibriSpeech, some of the models with the highest benchmark scores typed the removed number anyway in roughly 30-40 percent of examples — and typed the right one.[1]
There is no route to that output through the audio; the route runs through remembering the text. The strongest competing explanation is that the model guesses a plausible number from the surrounding sentences and now and then lands on the right one. The distribution in the study strains that reading: recovery rates are highest on the public benchmarks and lower on ep-fresh and libri-fresh, audio freshly collected from the same domains. Guessing from context would not swing that sharply from set to set.[1]
The spelling the benchmark is used to
The other two probes point the same way. The first uses an ensemble of low phoneme-error-rate models to flag the reference transcript's own mistakes: potential reference errors appear in 40 percent of the VoxPopuli clips analysed, which works out at roughly 3 percent of all reference words. Models with benchmark-fitted behaviour reproduce those faulty transcripts 18-30 percent of the time; in one example 6 of the 11 models drop a form of address the audio plainly contains, exactly where the reference transcript drops it. The third probe compares spellings that sound identical: several models clear the 50 percent random-choice baseline, and some reach about 90 percent accuracy at picking the convention their own benchmark is used to.[1]
To clear the random-choice baseline, a model first has to work out which set the audio came from. Once it can, part of the reported word error rate starts measuring how well it matches that set's spelling habits, and its link to transcription quality loosens. The gap may look small, but anyone choosing a model on that single number is choosing without seeing how it behaves outside the set.[1]
What a score can carry
The authors leave benchmark builders a concrete suggestion: replace independent and identically distributed splits with temporal, speaker or other metadata-based separation. If that separation is not adopted, I expect the same models to keep falling short of their own benchmark scores on freshly collected audio. The place to test it is settled too: the study added a "Benchmark fitting" tab to the Open ASR Leaderboard and open-sourced the analysis scripts. Per-model results published in that tab by 31 December 2026 will be enough either to bear the expectation out or to break it.[1]
In my 4 August column I wrote that when a generator's own detector reports no threshold, error rate or appeal path, what comes out is a statement of resemblance rather than a determination of origin. The gap here is of the same kind and sits one step earlier: without a control for set membership, a benchmark score describes the set along with the model. The encouraging part is that what would close it has a name this time — fully held-out evaluation sets, and fitting analyses published per model.[1], [2]