Eigen RadarAI
Analysis

Guardrail approvals expire before use as Claude searches its own alignment fixes

A 29 August arXiv preprint finds model guardrail approvals can be valid at check time and invalid at use, with verdict-change rates from 5.3 per cent to 48.4 per cent across five environments at eight simulator steps. Anthropic on 28 August reported Claude searching literature and training itself across 10 alignment failure categories, scoring 20 per cent better than human proposals on deception by its own measure.

Artificial Intelligence··Night
A matte gray robot hand extends a translucent card toward a blank reader at a closed turnstile in a bright glass-and-stone lobby; the card breaks apart before contact.

Guardrail approvals can fail between check time and use

In a self-adaptive system a model-based guardrail can issue an approval that is correct at check time yet no longer valid by the time it is acted on. A preprint on arXiv dated 29 August measures that gap: across five reproducible environments, at a common shift of eight simulator steps, the all-candidate verdict-change rate runs from 5.3 per cent to 48.4 per cent. The authors propose a shield that estimates each approval's validity horizon from its safe-side margin and recent volatility, without an explicit model of the plant's dynamics. At the same shift, oracle-labelled approval expiry falls from a 3.4-24.7 per cent range to a 0-1.8 per cent range. An audit of four judge models finds use-time invalidity above zero in every approval stream. The work is a preprint that has not been peer reviewed.[1]

Claude iterates alignment fixes across 10 failure categories

In research published on 28 August, Anthropic handed Claude 10 categories of alignment failure one at a time; each round the system searches the literature, proposes methods and data, trains the model and tests it. On deception, by the company's own measurement, Claude's best method scored 20 percent better than the best human proposal; 6 safety researchers closed 20 percent of the gap to a perfect score on average, while Claude reached 85 percent across multiple runs.[2]

Runtime guard timing and automated alignment search expose verification limits

The guardrail preprint reports verdict-change rates up to 48.4 per cent when approvals age eight simulator steps. Anthropic's automated alignment loop reports Claude reaching 85 percent on deception fixes by the company's own metric across multiple runs. The guardrail study is a preprint without peer review; Anthropic's scores are self-measured without independent audit in the pool.[1], [2]

References

  1. News sourcearXivThe approval was right when it was given and wrong when it was used↩1↩2
  2. News sourceAnthropicClaude looked for the alignment fix itself and outscored six safety researchers↩1↩2