Claude researches its own alignment fixes as Co-Scientist takes experiments into the lab
In research published on 28 August, Anthropic said Claude searched for methods across ten alignment failure categories, trained the model and on deception outscored six safety researchers' best proposals. Google DeepMind extended Co-Scientist into a closed-loop pipeline from hypothesis through lab equipment to manuscript drafts. Anthropic is automating alignment research while Google DeepMind expanded Co-Scientist into an experimental pipeline.
Artificial Intelligence··Morning
Claude searches for its own alignment methods
In research published on 28 August, Anthropic handed Claude ten alignment failure categories one at a time; each round the system searched the literature, proposed methods and data, trained the model and tested it. On deception, by the company's own measurement, Claude's best method scored 20 percent better than the best human proposal. Six safety researchers closed 20 percent of the gap to a perfect score on average, while Claude reached 85 percent across multiple runs. For a production model the system found more than 50 solutions in 60 hours and needed just over 2,000 training examples, which Anthropic puts at roughly 15,000 times more efficient than its own production alignment procedure.[1]
Anthropic lists the limits plainly
Anthropic lists the limits plainly: the failures studied were narrow next to production ones and political bias was not measured. Some failures are too rare or too recent for a benchmark to exist. Rejected methods were checked against only a limited set of predetermined capabilities. Evaluations such as Petri count as proxies for real-world misalignment, and whether the gains survive extensive reinforcement learning on other tasks was not tested.[1]
Co-Scientist also runs the experiment and writes the paper
Google DeepMind extended Co-Scientist, introduced in February 2025 on Gemini 2.0 as a hypothesis generator, into a closed-loop research system: it derives hypotheses, builds experimental plans, writes code, drives lab equipment, analyses results and drafts manuscripts. In materials science it designed synthesis recipes for two-dimensional materials; semiconductor thin films were ready on the first attempt. In biology the system built an autonomous image analysis pipeline that predicts E. coli colony patterns with 75 percent accuracy on shape features. In computer science it designed a medical AI architecture called Agent_H without human intervention; that design's results did not hold up under physician evaluation. With the reliability modules switched on the fabrication rate fell to 4 percent, against 46 percent measured with them off.[2]
Related columns
For more information on this topic, you can read the related columns.