Agents shrink the dataset and swap a component for a lookup, and the result still comes out as a finding
A preprint published on arXiv on 29 August names systematic deviations in scientific reproduction — shrinking datasets, replacing components with lookup tables, generalising from limited resources — "methodological hallucinations"; the proposed ABE-Ralph framework reports a 93 percent robust execution rate across 30 long-horizon runs. THE DECODER reports Google DeepMind's Co-Scientist medical architecture design did not hold up under physician evaluation and that fabrication reached 46 percent with reliability modules off. The two developments complement same-week warnings about method discipline in autonomous research.
Artificial Intelligence··Evening
ABE-Ralph names methodological hallucinations
A preprint published on arXiv on 29 August argues that an agent reproducing a scientific result has to do more than produce code that runs: it has to implement the reference method faithfully, design an experiment that tests the paper's claim, and supply evidence supporting it. The authors call the systematic failures at those three steps "methodological hallucinations"; listed failures include reducing the dataset, replacing a component with a lookup function, and drawing conclusions from resource-limited settings. The proposed ABE-Ralph framework represents claims, protocols, required components, baselines and metrics as structured experimental constraints. The authors report a 93 percent robust execution rate across 30 long-horizon runs spanning 12 machine learning domains, and performance matching or exceeding the state of the art on 5 of 23 NatureBench discovery tasks. The work is a preprint, and the rates come from the authors' own evaluation rather than independent replication or peer review.[1]
Co-Scientist delivered a design that failed physician review
According to THE DECODER, Google DeepMind extended Co-Scientist into a closed loop from experiment planning to lab equipment, designing synthesis recipes and obtaining semiconductor thin films on the first attempt. In biology it built a pipeline predicting E. coli colony patterns with 75 percent accuracy. In computer science it designed a medical AI architecture called Agent_H without human intervention; that design's results did not hold up under physician evaluation. With reliability modules on the fabrication rate fell to 4 percent, against 46 percent measured with them off. THE DECODER reports the system delivered concrete outputs in biology and materials science while failing on the medical architecture; every figure is the authors' own measurement. THE DECODER also notes the February 2025 hypothesis-generator framing has expanded into a loop that now writes code and drives equipment, with no independent replication, peer review or outside audit reported for the new numbers.[2]
Method-discipline warnings arrived two ways in the same week
ABE-Ralph names deviations such as shrinking datasets in reproduction and claims a 93 percent robust execution rate, while on the Co-Scientist side fabrication reached 46 percent with modules off. One shows tightening method with preventive constraints, the other shows reliability modules cutting fabrication in a lab loop. ABE-Ralph results come from the authors' own runs across 30 long-horizon episodes in 12 machine learning domains; on Co-Scientist one architecture also failed physician evaluation. ABE-Ralph reports matching or beating the state of the art on 5 of 23 NatureBench tasks; Co-Scientist reports 75 percent accuracy in biology. For readers the concrete development is an academic framework and a field measurement on autonomous research method reliability arriving in the same week without peer review, with neither line yet reporting independent replication, peer review, outside validation, independent audit or third-party confirmation of the other's headline numbers.[1], [2]