The design
Abraham Camelo-Guerrero and Jairo Diaz-Rodriguez compared reviews written by OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview and Anthropic Claude Opus 4.6 with human assessments of 300 topic-matched ICLR 2026 submissions, split equally across oral, poster and rejected papers. Each model reviewed the same papers under standardised instructions with the decision information removed.[1]
Holding the papers, the prompt and the hidden decision constant across three providers is what makes the comparison readable. It is also what bounds it: the sample is one venue, one year and one submission population, so what the study can support is a statement about these three models on ICLR 2026, not about reviewing in general.[1]
Where the agreement stops
All three models separated accepted from rejected papers. None reproduced the oral versus poster distinction present in the human ratings. Those two sentences describe different tasks: the first is a coarse threshold that a broad quality signal can clear, the second requires ranking within the accepted set, where the differences are smaller and the human decision partly concerns programme composition as well as the paper.[1]
Scoring differed by provider. Gemini gave consistently higher ratings, while the OpenAI and Anthropic models tracked human judgements more closely on rejected and poster papers but were too harsh on oral papers. A provider-level offset of that kind is the sort of thing calibration can absorb. Being wrong specifically on the strongest papers survives calibration, because the fine distinction is about the papers at the top of the distribution.[1]
What the reviews were about
The content difference is more informative than the scores. The models flagged missing baseline comparisons more often than humans did, and humans raised computational-efficiency concerns more often. Two reviewing processes can therefore agree on an outcome while disagreeing about what a paper's weakness is, and an agreement statistic computed on decisions will not show that.[1]
The authors say as much: broad decision alignment does not imply agreement with finer human judgements or reviewing priorities. That is the conclusion the design supports, and it should not be upgraded. The work is a preprint and has not been peer reviewed, which is worth stating twice for a paper about peer review. The measurement that would extend it is a second venue with a different acceptance rate, where the oral and poster boundary is drawn somewhere else.[1]