Eigen RadarAI
Analysis

A safety check that resets each turn misses an attack split across turns

LoopHarness, published on arXiv on 29 August, shows mathematically that safeguards resetting after each iteration fail when evidence is split across turns in autonomous agent loops. The same day's SARA preprint separates the moment a tool's output starts naming an action and reports attack success below 0.63 percent. Kotaku's account of the Cutting Room Floor attack describes a small archive going offline after a language-model ban; the site holds a banned user responsible.

Artificial Intelligence··Evening
At a bright, unbranded security checkpoint, the scanner lamp lights only a cane's curved handle in the current tray while its straight lower section sits in the next tray.

LoopHarness shows attacks split across turns

A preprint published on arXiv on 29 August says that in autonomous agent loops the usual safeguards reset after each iteration, and they fail when the evidence is fragmented across several iterations. The authors show mathematically that trajectory-scoped monitors cannot separate true positives from false positives against such an attack. Geometric decay of the risk score is not enough either: a constant cooling-off period can be exploited regardless of horizon length. The proposed LoopHarness keeps safety state persistent and non-decaying at the loop level, and under stated conditions bounds the expected number of unauthorised irreversible actions by a constant independent of the horizon length N. Evaluation protocols, ablations and adaptive red-team testing against Agent-SafetyBench tasks are also presented. The work is a preprint without peer review and addresses multi-turn attacks that per-turn checks miss in production-style agent loops, with no independent replication reported.[1]

SARA separates action suggestion from authorisation

A preprint posted to arXiv on 29 August treats the moment a tool's output stops carrying data and starts naming an action as the point where an agent can be steered. The authors propose SARA, which separates the question of what suggests an action from the question of what is authorised to run it, and holds authorisation to the user's stated goal and audited evidence. A context-isolated Action Probe exposes the action-inducing part of a tool's output, while a rule the authors call No-History-Promotion stops earlier conversation from turning into authority to act. On the AgentDojo and AgentDyn benchmarks they report attack success no higher than 0.63 percent across four main settings. The reported figures come from the authors' own runs on two public benchmark suites, and the work is a preprint that has not been peer-reviewed or independently replicated by outside teams or field auditors.[2]

The game archive closed after a language-model block

The Cutting Room Floor, a game archive founded in 2002, went offline after a denial-of-service attack. The site had been serving its own MS Paint images to requests coming from language models, and it holds responsible a user it had banned for reaching the site through Claude. By the site's account, the banned user wrote on X that he would complain to the hosting provider and look into the domain and its operators, then evaded the ban and filed abuse reports. Co-founder Xkeeper said he had made no direct accusation and had only noted that the post came immediately before the attack. That is the site's own reading, and it carries no independent confirmation. The archive had tried to slow automated access by returning deliberately low-quality images; the attack ended in infrastructure outage for a small volunteer project that had already been blocking model traffic before going offline.[3]

References

  1. News sourcearXivA safety check that resets each turn misses an attack split across turns↩
  2. News sourcearXivWhen a tool's output starts giving orders, the agent needs a second gate↩
  3. News sourceKotakuA game-archive wiki went offline after an attack that followed its language-model ban↩