OWASP study finds expert rankings diverge from incident data
OWASP researchers compared security risks in large-language-model applications with 7,714 incidents. Prompt injection ranked first with experts but twelfth in incidents, while misinformation ranked thirteenth for experts and second in incidents. Agreement was measured at 0.20. The study labelled 6,639 incidents across 20 categories, and changing classifier precision adds uncertainty to the resulting distribution.
Artificial Intelligence··Night
The two rankings do not produce the same order
Two leaders of the OWASP Top 10 for LLM Applications compared the expert ranking with a dataset of 7,714 security incidents. Prompt injection ranked first for experts but twelfth in incidents, while misinformation ranked thirteenth for experts and second in incidents.[1]
Agreement was measured as low
Agreement between the lists, measured with Cohen's kappa, was 0.20. The researchers say this shows that expert prioritisation and the distribution of documented incidents do not send the same signal; the comparison alone does not establish that one list invalidates the other.[1]
Classification uncertainty carries into the results
The study labelled 6,639 incidents against a 20-entry taxonomy and applied a Bayesian correction for classifier error. Classifier precision varied widely across categories. The incident distribution is therefore a counted observation, but it cannot be read independently of measurement uncertainty. The comparison's value lies in placing expert lists and documented incidents on the same scale. But because incident classification has variable precision, the ranking gap must be assessed in the context of both security priorities and how the data was labelled. OWASP researchers compared security risks in large-language-model applications with 7,714 incidents. Prompt injection ranked first with experts but twelfth in incidents, while the measurement method also makes the uncertainty in the results explicit. This briefing keeps together the actor, reported scope and stated uncertainties of the event described by the source. It adds no separate development, definite outcome or comparison absent from that source; the headline isolates only this event.[1]