The link between the proxy and the target

A bug bounty programme uses the number of incoming reports as a proxy for something it cannot measure directly: the number of genuine vulnerabilities found. The proxy works only while writing a report costs a certain amount of effort. According to a report citing the Financial Times, Apple has limited how many reports researchers may submit and introduced a 30-day waiting period after a flood of low-quality AI-generated submissions containing hallucinated flaws clogged its review pipeline. When the cost of producing a report falls, the composition of the counted set changes and the proxy stops tracking the target.[1]

Where the intervention is applied matters. The cap and the waiting period attach to the person submitting rather than to the quality of the report. A threshold of that kind behaves like a classifier with no discriminating power: it lowers the acceptance rate, but decides who ends up on which side by submission history. A filter that runs on submitter identity rather than on a quality criterion screens out correct reports at the same rate as wrong ones.[1]

The two-sided cost of the threshold

One instance of that cost is documented. The Italian startup Bynario used ChatGPT to find a serious macOS vulnerability that could give an attacker full control of a machine, but could not report it because Apple had blocked further submissions; chief executive Alfredo Pesoli estimates the flaw's black-market value at $100,000 to $200,000. Apple has since reached out to Bynario, according to the report. A single case gives no rate, but it does show that the threshold can produce a false negative, a cost invisible to any assessment that looks only at the acceptance rate.[1]

The numbers needed to assess the threshold are missing. The report does not give the date the cap took effect, how many submissions were screened out on quality grounds, or the criterion by which a request for a higher quota is decided. The same report says Apple uses models from Anthropic and OpenAI in its own vulnerability hunting and that its latest updates carried five times as many fixes as usual. That too is a volume figure; without the severity of the fixed flaws and their likelihood of exploitation it says nothing about how close the programme is to its target. The minimum dataset that would answer the question is narrow: the effective date, the number of screened-out submissions, and the outcome of quota appeals.[1]