Eigen RadarAI
Analysis

Only 6.52 per cent of neuro-symbolic studies rerun as Co-Scientist widened its lab loop

An audit by nine authors published on arXiv on 29 August fully or partially reproduced only 85 of 1,304 eligible neuro-symbolic studies, or 6.52 percent. Google DeepMind extended Co-Scientist into a closed-loop pipeline from hypothesis through lab equipment to manuscript drafts. Automation is accelerating while most published artifacts still cannot be rerun.

Artificial Intelligence··Midday
On a long sunlit library table, irregular stacks of closed notebooks recede behind a few blank open notebooks and a working brass mechanical device.

The audit screened thousands of bibliography entries

An audit by nine authors, published on arXiv on 29 August, worked through the neuro-symbolic AI literature in six stages: 5,497 bibliography entries collected, 3,018 duplicates removed, 1,304 studies left eligible. The team fully or partially reproduced 85 of them, or 6.52 percent. Of the attempted reruns, 321 stopped with non-code material missing and 42 with the code repository missing or unusable. The authors describe a lasting reproducibility deficit in the field and argue that publication should require complete, versioned and permanently archived artifact bundles. They offer their six-stage framework as a reusable procedure rather than a one-off count. The work is a preprint on arXiv that has not been peer-reviewed, and the audit treats missing non-code artifacts as the largest single stop reason among attempted reruns. The reported findings come from the authors' own runs, and the work is presented as a preprint that has not been peer-reviewed.[1]

Co-Scientist moved the experimental line into a closed loop

Google DeepMind extended Co-Scientist, introduced in February 2025 on Gemini 2.0 as a hypothesis generator, into a closed-loop research system: it derives hypotheses, builds experimental plans, writes code, drives lab equipment, analyses results and drafts manuscripts. In materials science it designed synthesis recipes for two-dimensional materials; semiconductor thin films were ready on the first attempt. In biology the system built an autonomous image analysis pipeline that predicts E. coli colony patterns with 75 percent accuracy on shape features. In computer science it designed a medical AI architecture called Agent_H without human intervention; that design's results did not hold up under physician evaluation. With the reliability modules switched on the fabrication rate fell to 4 percent, against 46 percent measured with them off. Every figure is the authors' own measurement, and the closed loop now covers equipment control and manuscript drafting rather than hypothesis generation alone. The reported findings come from the authors' own runs, and the work is presented as a preprint that has not been peer-reviewed.[2]

Faster automation meets older artifact gaps

The audit quantifies how most published neuro-symbolic work could not be rerun for missing code or non-code artifacts, a reminder that new output from automated research lines may face similar archiving pressure. On the Co-Scientist side a medical architecture that failed physician evaluation is also reported; every figure is the running team's own measurement. The two reports describe opposite ends: one the past literature that cannot be reproduced at scale, the other a claim to automate future experiments through lab equipment and drafts. Co-Scientist reports first-attempt thin-film success in materials science. For readers the shared frame is tension between the speed of AI-assisted research and the portability of published evidence across both symbolic and automated pipelines. The reported findings come from the authors' own runs, and the work is presented as a preprint that has not been peer-reviewed.[1], [2]

References

  1. News sourcearXivOnly 85 of 1,304 neuro-symbolic studies could be rerun from what they published↩1↩2
  2. News sourceTHE DECODERCo-Scientist now plans the experiment, runs the lab equipment and writes the paper↩1↩2