Eigen RadarAI
Analysis

Dressed-up market data pushed agents into calls they could not support

In a preprint published on arXiv on 29 August, 12 frontier models were shown professional-looking market data on questions no data could settle; as the evidence looked stronger, commitment rose from 6.5 percent to 54.0 percent of runs. Fabricated numbers produced a 36.8 percent commitment rate against 37.6 percent for genuine data. Anthropic's automated alignment research announced on 28 August also outscored human proposals on deception.

Artificial Intelligence··Midday
Three uneven stacks of blank paper sit on a walnut desk in golden-hour light; a magnifying glass reveals fine fibres unravelling from one sheet's corner.

Stronger-looking evidence raised commitment

In a preprint published on arXiv on 29 August, 12 frontier models were shown professional-looking market data on questions no data could settle. As the evidence was made to look stronger, the share of runs in which a model committed to an answer rose from 6.5 percent to 54.0 percent. When every displayed number was fabricated the commitment rate was 36.8 percent, against 37.6 percent for genuine data. The author does not place the failure in the models' sense of what is knowable but in the step that decides whether to act: asked in advance to classify the same questions, the models identified the unanswerable ones about 90 percent of the time. Fine-tuning a small model removed the behaviour on the original test cases and carried over to domains it had not seen, although the effect did not survive response formats that block reasoning. The study is a preprint on arXiv and has not been peer-reviewed.[1]

Anthropic announced automated search on deception

In research published on 28 August, Anthropic handed Claude 10 alignment failure categories one at a time; each round the system searched the literature, proposed methods and data, trained the model and tested it. On deception, by the company's own measurement, Claude's best method scored 20 percent better than the best human proposal; six safety researchers closed 20 percent of the gap to a perfect score on average, while Claude reached 85 percent across multiple runs. For a production model the system found more than 50 solutions in 60 hours and needed just over 2,000 training examples, which Anthropic puts at roughly 15,000 times more efficient than its own production alignment procedure. Anthropic lists plainly that the failures studied were narrow next to production ones and political bias was not measured; some failures are too rare or too recent for a benchmark to exist. Evaluations such as Petri count as proxies for real-world misalignment, and whether the gains survive extensive reinforcement learning on other tasks was not tested.[2]

Two lines seek intervention at different stages

The preprint shows an agent can be swayed by the surface of evidence so that commitment rises whether or not the numbers are real, a direct measure of how tables in financial interfaces can move the action threshold. Anthropic's automated research aims to speed the production alignment loop by searching for methods and training on the deception category; rejected methods were checked against only a limited set of predetermined capabilities. The two reports share a risk vocabulary but focus on data presentation at decision time and on post-training measurement respectively. For readers the shared message is that agents do not stay silent in the face of plausible-looking inputs; where intervention should sit depends on which line of work is in view, and neither study claims to cover political bias or full production failure modes.[1], [2]

References

  1. News sourcearXivDressed-up market data pushed AI agents into calls they could not support↩1↩2
  2. News sourceAnthropicClaude looked for the alignment fix itself and outscored six safety researchers↩1↩2