Dress the panel and commitment multiplies
In a preprint Aggarwal published on arXiv, 12 frontier models were shown professional-looking market panels on questions no dataset could settle. As the evidence was made to look stronger, the share of runs in which a model committed to a directional call rose from 6.5 per cent to 54.0 per cent. When every figure on the panel was invented—so nothing the model could see was true except the question itself—commitment climbed from 24.5 per cent to 36.8 per cent, statistically indistinguishable from the 37.6 per cent produced by genuine market data. The authority of the packaging appears to unlock action.[1]
On matched answerable questions attached to the same panels, the same models answered essentially always, at near-perfect accuracy—incapacity does not explain the pattern. Stated probabilities barely moved across the gradient that swung action by 48 points. The numbers on screen may be shifting the act decision.[1]
The knowing gate and the acting hand are separate
Asked to classify knowability before acting, models called questions irreducible 90 per cent of the time and then committed on just 0.4 per cent of those. The gate looks narrow and functional—until a panel intervenes. Visual evidence placed in front of an agent may bypass the gate that blocks action without answering the question.[1]
SARA, described in a separate arXiv preprint, splits observation from execution authority: when a tool output stops carrying data and starts naming an action, a context-isolated Action Probe tracks where the action came from, and actual tool calls are authorized only against the user's objective and audited evidence. Across AgentDojo and AgentDyn, attack success stayed no higher than 0.63 per cent in four primary settings. The two papers test different surfaces—market panels versus tool outputs—but both try to stop an untrusted observation from becoming permission to act.[1], [2]
A brake you can train but not set once
Aggarwal shows the gate is separable and trainable: supervised fine-tuning of a 3B model on 540 synthetic cases drove commitment to 0.0 per cent on the original cases and transferred to three unseen domains. Rigid response formats that removed room to reason left the model confident and wrong—the gate is trainable and context-fragile, and deployment needs both halves of that sentence. Whenever an institution feeds an agent a panel, a tool output, or any other evidence wrapper, checking which format leaves the gate engaged is now a question separate from the model card.[1]