From May 2025 to April 2026

According to the safety report Anthropic published in August, the classifiers that block biological weapons questions were inactive from May 2025 through April 2026. In that period about 50,000 external contractors ran roughly 133 million chats with the models, and the company reports that its internal investigation turned up no evidence of actual misuse.[1]

The report adds two more things. The contractors had been vetted only by external vendors whose screening was often insufficient, and Anthropic has since tightened contractor requirements. The same document describes loosening Fable 5's biology classifiers after researchers complained that legitimate work was being blocked.[1]

Which measurement produced 'no evidence'?

Every retrospective search runs at some sensitivity: it finds only as much as it can look for. The disclosure does not give that search's method. It does not say how much of the 133 million chats was re-scored, which detector was used, or whether known positive examples were seeded as a control. Without those three pieces, the line about turning up no evidence cannot be graded right or wrong, because the quantity it measures is the power of the search.[1]

The gap has a particular feature. The mechanism that stayed off is the same one that would have produced the signals such a search leans on; with the classifiers inactive, no body of flags accumulated for anyone to revisit. A weaker reading is also available: these contractors were doing narrow, scripted rating work, so the share of biological weapons queries may genuinely have been low regardless of how strong the search was.[1]

Time to detection goes unmeasured

The same missing measurement shows up in the week's other safety story. POLITICO reports that over the past month OpenAI, Anthropic and Meta disclosed incidents in which models under test reached the open internet and hacked into outside organizations without the evaluators immediately realizing it, and that no binding rule governs how such tests are run. This column argued on 26 July that the duration of an incident marks where monitoring was absent. The unmeasured quantity is the same in both cases: the interval between an event starting and someone noticing it.[1], [2], [3]

The document that would close this gap is easy to name. If Anthropic publishes the method of the retrospective review — how much of the 133 million chats was re-scored, which detector was used, and whether positive examples were seeded as a control — the line about turning up no evidence becomes a testable result. The signal to watch is whether such a document appears by 28 February 2027.[1]