The loop closed; the benchmark list did not
In the research Anthropic published on 28 August, Claude runs the steps of alignment research itself: it searches the literature, proposes methods and data, trains the model, then tests it. The company defined 10 categories of alignment failure and used 3 to 5 benchmarks for each. On deception, Claude's best method scored 20 percent better than the best human proposal; 6 safety researchers closed 20 percent of the gap to a perfect score on average, while Claude reached 85 percent across multiple runs. For a production model, over 50 solutions came out in 60 hours, and just over 2,000 training examples were enough.[1]
The layer that produces this result is the search loop built around the model. The same Claude does not reach a stronger result when it is pointed in a direction a human researcher proposed; the gain comes from training and discarding many methods quickly. Another reading is possible: if the benchmarks resemble one another closely enough, what the loop finds may be a fix suited to that benchmark family rather than a general alignment method. Anthropic's own list of limits feeds that second reading, because rejected methods were checked against only a limited set of predetermined capabilities.[1]
The check that comes from outside
Google DeepMind's Co-Scientist takes the same loop into the lab: it derives hypotheses, builds experimental plans, writes code, drives instruments and drafts manuscripts. In materials science it designed the synthesis recipe for two-dimensional materials and produced semiconductor thin films on the first attempt. In computer science it designed a medical architecture called Agent_H without human intervention, and that design's results did not hold up under physician evaluation. With the reliability modules on, the fabrication rate was measured at 4 percent; with them off, 46 percent.[2]
The constraint the two systems share is that their result is only as solid as a check from outside. Anthropic writes for itself that the failures studied were narrow next to production ones, that political bias was not measured, and that some failures are too rare or too recent for a benchmark to exist; it also counts evaluations such as Petri as proxies for real misalignment. On the Co-Scientist side the check did its work concretely: physician evaluation did not confirm the results of the architecture the system designed without human intervention. In both cases the verdict comes from the measurement standing outside the search loop.[1], [2]
What changed for the builder
On this desk on 27 August I wrote that evaluating agents is leaving the SDK and moving into the tracing layer; the automated alignment researcher shows the same move one layer up. As searching for a method gets cheap, the asset that has to be maintained moves from the method to the benchmark set. Anthropic's 10 categories, with 3 to 5 benchmarks each, are the name of an inventory an organisation has to keep, and the questions of who writes it, who updates it and how its coverage is measured arrive with it.[1], [3]
The measurable signal is this: an independent team that runs the same loop on a failure benchmark absent from Anthropic's list, and reports a comparable gain, moves the result off a fit to that benchmark family. If no such run is published by 30 November 2026, what remains stands as the company's own measurement. Anthropic already writes that it did not test whether the gains survive extensive reinforcement learning on other tasks, and the party able to run that test is the party holding the benchmark set.[1]