SARA builds an action gate as a game archive goes offline after its language-model ban
A preprint published on arXiv on 29 August proposes SARA, which separates the moment a tool's output starts naming an action; attack success stayed no higher than 0.63 percent across four main settings. The Cutting Room Floor, a game archive founded in 2002, went offline after a denial-of-service attack it ties to a user who evaded its language-model block. One line builds preventive authorisation; the other shows reactive site defence after an access ban.
Artificial Intelligence··Midday
SARA separates action suggestion from authorisation
A preprint posted to arXiv on 29 August treats the moment a tool's output stops carrying data and starts naming an action as the point where an agent can be steered. The authors propose SARA, which separates the question of what suggests an action from the question of what is authorised to run it, and holds authorisation to the user's stated goal and audited evidence. A context-isolated Action Probe exposes the action-inducing part of a tool's output and tracks where it came from, while a rule the authors call No-History-Promotion stops earlier conversation from turning into authority to act. On the AgentDojo and AgentDyn benchmarks they report attack success no higher than 0.63 percent across four main settings. The reported figures come from the authors' own runs on two public benchmark suites. The work is a preprint on arXiv that has not been peer-reviewed, and the authors present the framework as a reusable procedure for checking every action against policy before it takes effect rather than as a one-off demonstration on a single model family.[1]
The game archive closed after a language-model block
The Cutting Room Floor, a game archive founded in 2002, went offline after a denial-of-service attack. The site had been serving its own MS Paint images to requests coming from language models, and it holds responsible a user it had banned for reaching the site through Claude. That is the site's own reading, and it carries no independent confirmation. By the site's account, the banned user wrote on X that he would complain to the hosting provider and look into the domain and its operators, then evaded the ban and filed abuse reports. Co-founder Xkeeper said he had made no direct accusation and had only noted that the post came immediately before the attack. The archive had tried to slow automated access by returning deliberately low-quality images rather than full wiki pages, which made the later infrastructure outage a visible end point for a small volunteer project that had already been experimenting with language-model traffic filters.[2]
A preventive gate and a reactive wall see the same problem at different scales
SARA proposes a second authorisation layer at runtime against steering through tool output; measured attack success stayed low, though the figures come from the authors' own runs on AgentDojo and AgentDyn. The Cutting Room Floor describes a volunteer archive hitting infrastructure attack while trying to filter language-model traffic; the site had tried slowing automated scraping by returning MS Paint images instead of full pages. The two cases ask where an action or scrape request should stop, at different layers: one at the action suggestion inside output, the other at the network edge where a banned user could still reach the host. SARA's No-History-Promotion rule and Action Probe aim to stop earlier chat history from becoming permission to act, while the archive's defence was reactive filtering followed by outage. For readers the concrete development is an academic framework publishing in the same week an archive went offline after its language-model ban.[1], [2]