Eigen RadarAI
Analysis

PHRBench separates correcting a false premise from finding the right answer

PHRBench follows how eighteen language models handle deliberately false context across chemistry, biomedicine, physics and coding. A correct answer can arrive without an explicit correction of the premise, so the study tracks both behavior and outcome. The single preprint finds accuracy falls under erroneous context, and correction attempts do not always succeed.

Artificial Intelligence··Evening
A red plug beside an open socket and two white test housings connected to separate cables on a yellow work surface.

False context lowers answer accuracy

Incorrect information from an earlier model can become the premise for a later answer. PHRBench, a controlled evaluation designed by Linghao Meng and colleagues, follows that process across eighteen language models. Its single preprint covers 4,820 instances in chemistry, biomedicine, physics and code generation. False context lowered accuracy throughout the evaluated group.[1]

Accepting, bypassing and correcting are measured separately

Responses are classified by whether they accept the false premise, avoid it or explicitly correct it. Successful recovery requires both a correction and a correct final answer. Larger models in several families attempted more corrections but also suffered greater accuracy degradation under erroneous context. A correction attempt could still end with an incorrect answer.[1]

Each base question receives truthful context and variants that alter one semantic element. The errors contradict a domain rule, distort relevant conditions or add an invalid scientific explanation. Multiple-choice and open-ended tasks are both included, with changes designed to avoid revealing the answer.[1]

Prompt features help estimate recovery

The team extracted twenty-four structural and semantic features to estimate recovery ahead of an answer. Error location, the length of the prompt and violated domain rules carried predictive information in this setup. These deliberately constructed tasks measure responses to specified false premises; they do not establish a deployment-wide hallucination rate.[1]

References

  1. News sourcearXivPHRBench tests how language models recover from false premises↩1↩2↩3↩4