Eigen RadarAI
Analysis

Twenty-six authors claim frontier-level gains on a continually trained open model as Claude searches its own alignment

The Thomson preprint published on arXiv on 29 August reports a continual-learning stack on open-weight models approaching today's frontier systems on agentic tasks. Anthropic's 28 August research says Claude searched 10 alignment failure categories on its own; on deception it reports a result 20 percent better than the best human proposal. THE DECODER writes Google DeepMind extended Co-Scientist into a closed loop that drives lab equipment. Three developments in the same week advance models' own research and improvement loops.

Artificial Intelligence··Evening
In a daylight research atrium, a robotic gantry adds a new segment to the bright outer edge of a large transparent lattice whose older, darker core remains intact.

Thomson claims continual learning on an open model

A preprint published on arXiv on 29 August describes Thomson, a continual-learning stack applied to open-weight models. The authors report results they place alongside current frontier systems on agentic tasks, safety, legal and tax questions, multilingual work and large-scale research, and say the gains arrive at compute and staffing budgets well below what is usually assumed. The stack combines mid-training and post-training steps with safeguards the authors say keep a model able to learn without losing what it already knows. They describe the outcome as a pattern in which capability rises across a wide range while forgetting is almost entirely removed. The claims are the authors' own; no independent replication is reported. The work is a preprint that has not been peer-reviewed, with twenty-six authors listed on the paper.[1]

Claude searched for its own alignment fix

In research published on 28 August, Anthropic handed Claude 10 categories of alignment failure one at a time; each round the system searches the literature, proposes methods and data, trains the model and tests it. On deception, by the company's own measurement, Claude's best method scored 20 percent better than the best human proposal; 6 safety researchers closed 20 percent of the gap to a perfect score on average, while Claude reached 85 percent across multiple runs. For a production model the system found over 50 solutions in 60 hours and needed just over 2,000 training examples, which Anthropic puts at roughly 15,000 times more efficient than its own production alignment procedure. The company lists limits plainly: the failures studied were narrow next to production ones and political bias was not measured, and whether the gains survive extensive reinforcement learning on other tasks was not tested in the published account.[2]

Co-Scientist moved from experiments to a lab-equipment loop

According to THE DECODER, Google DeepMind extended Co-Scientist as an autonomous research pipeline: in materials science it designed synthesis recipes and obtained semiconductor thin films on the first attempt. In biology it built a pipeline predicting E. coli colony patterns with 75 percent accuracy. In computer science it designed a medical AI architecture called Agent_H without human intervention; that design's results did not hold up under physician evaluation. With reliability modules on the fabrication rate fell to 4 percent, against 46 percent measured with them off. THE DECODER reports the February 2025 Gemini 2.0 hypothesis-generator launch has expanded into a loop that now writes code and drives equipment; every figure comes from the team's own measurement and no independent replication, peer review, outside validation or independent audit is reported for the expanded loop.[3]

References

  1. News sourcearXivTwenty-six authors put frontier-level claims on a continually trained open model↩
  2. News sourceAnthropicClaude looked for the alignment fix itself and outscored six safety researchers↩
  3. News sourceTHE DECODERCo-Scientist now plans the experiment, runs the lab equipment and writes the paper↩