Eigen RadarAI
Analysis

PatchBench catches safety fixes that block harmless questions

PatchBench tests whether a repair stops harmful answers while preserving nearby harmless requests. Its new preprint builds on verified failures from eight language models and separates those two outcomes. Some tested interventions leave broad capability scores nearly unchanged while sharply reducing responses to benign questions with similar wording or structure.

Artificial Intelligence··Evening
On a bright testing bench, red and teal spheres in two metal tracks stop before a wide black barrier spanning both paths.

Harmless neighboring questions lose answers

Safety interventions can preserve a language model’s overall capability score while blocking harmless questions near the targeted harmful request. PatchBench, a test collection for repairing observed harmful behavior, examines that local damage. The single preprint compares four methods that steer a model’s internal activations. Some interventions on Llama and Gemma leave the broad knowledge test MMLU almost unchanged while sharply reducing benign-neighbor preservation.[1]

Verified harmful answers define the repair targets

The researchers began with 27,870 prompts from 37 public datasets. After screening, they queried eight models with 15,314 English prompts. The resulting set contains 400 verified failures after automated filtering, ranking and human review, with fifty for each model. The subsets differ because a question that breaks one model can be refused by another.[1]

Each failure gets harmful and benign variants

PatchBench-Local creates twenty-nine neighbors per original prompt: nineteen retain harmful intent, while ten use benign intent with matching structure or vocabulary. Harmful-request correction and benign-response preservation receive separate scores. The collection is distributed for defensive research under a non-commercial license, with intended-use documentation and dual-use warnings. Its coverage does not certify safety against every attack.[1]

References

  1. News sourcearXivPatchBench tests the collateral damage of safety patches↩1↩2↩3