Eigen RadarAI
Analysis

O*NET-BENCH separates judges’ ranking accuracy from acceptance estimates

The O*NET-BENCH preprint compares AI judges with workers’ ratings of responses to occupational tasks. Across 33 judge configurations, better agreement on response ordering can accompany worse agreement on average scores and acceptance rates. Calibration reduces mean bias but leaves individual disagreements. The audit uses existing survey data and treats worker ratings as target judgments rather than objective truth.

Artificial Intelligence··Night
An assessor examines the metal bracket joining a tabletop timber beam and column model.

Ordering and acceptance are separate measures

The O*NET-BENCH preprint finds that language-model judges can agree more closely with workers on the ordering of AI responses while matching their average scores or acceptance rates less well. In one fine-tuned lineage, presentation, examples and output format changed together. The study does not isolate presentation’s effect from the other changes. It evaluates response ordering separately from mean scores and the share judged acceptable.[1]

An existing worker survey supplies the data

The audit in this single preprint reuses 45,796 ratings gathered in an earlier survey, with no additional human data collected. US participants recruited through Prolific assessed tasks in occupations where they reported experience. Ratings run from one to nine, with seven or higher classed as acceptable. The test split contains 4,501 ratings, against 33 judging configurations drawn from six families of models. Workers and task identifiers are separated across the data splits.[1]

Calibration leaves individual disagreements

Calibration largely corrects the targeted bias in the average, but substantial disagreements on individual ratings remain. Each response has only one worker rating, and the audit cannot establish whether the disagreement reflects the worker’s rating or the model’s. Worker judgments are the study’s comparison target, not assumed objective truth. Some test analyses are exploratory, and response order in listwise prompts is not randomized. Findings concern these evaluation protocols and do not measure job displacement.[1]

References

  1. News sourcearXivO*NET-BENCH finds ranking agreement can mask acceptance-rate errors↩1↩2↩3