Eigen RadarAI
Analysis

Agents knocked real controllers off target in 75 of 240 runs

PLCBench, published on arXiv on 29 August, measured language-model agents sustaining a physical objective on programmable logic controllers in 31.3 percent of runs; 98 episodes failed before basic reads succeeded. THE DECODER reports Google DeepMind extended Co-Scientist into a closed loop from experiment planning to lab equipment, designing synthesis recipes and obtaining semiconductor thin films on the first attempt. One case shows industrial process risk, the other lab automation in the same week.

Artificial Intelligence··Evening
On a bright factory floor, a robot arm works above unlabelled containers bunched diagonally against a guide rail as a back-turned operator reaches for a manual lever beside an amber warning beacon.

PLCBench measured 31.3 percent success on real controllers

PLCBench, published on arXiv on 29 August, measures whether language-model agents given access to programmable logic controllers can produce sustained physical impact, using a hardware-in-the-loop setup. Across five model families and 240 episodes on real controllers, 75 episodes — 31.3 percent — sustained their physical objective. Broken down by stage, 98 episodes failed before basic system reads succeeded, while 62 reached process-linked writes without hitting the final objective. When the agents had better sensor data, conditional objective success rose from 44.2 percent to 64.0 percent. The authors present the framework as a map of defensive intervention points, making it measurable which stage is worth blocking. The work is a preprint without peer review, and the numbers belong to a laboratory rig rather than a live plant. The paper frames the measurement as answering where to stop an agent before physical impact spreads across a controlled industrial interface.[1]

Co-Scientist moved from experiment plans to equipment

According to THE DECODER, Google DeepMind has extended Co-Scientist, introduced in February 2025 on Gemini 2.0 as a hypothesis generator, into a closed-loop research system: it derives hypotheses, builds experimental plans, writes code, drives lab equipment, analyses results and drafts manuscripts. In materials science it designed synthesis recipes for two-dimensional materials and obtained semiconductor thin films on the first attempt. In biology the system built an autonomous image analysis pipeline that predicts E. coli colony patterns with 75 percent accuracy on shape features. In computer science it designed a medical AI architecture called Agent_H without human intervention; that design's results did not hold up under physician evaluation. With reliability modules on the fabrication rate fell to 4 percent, against 46 percent measured with them off. THE DECODER reports the February 2025 hypothesis-generator framing has given way to a loop that now writes code and drives equipment, with every figure coming from the team's own measurement rather than independent replication.[2]

Factory controllers and a lab loop landed in the same week

PLCBench reports a 31.3 percent sustain rate on industrial controllers while Co-Scientist describes driving lab equipment successfully on a first attempt; both measure agents reaching into the physical world. PLCBench numbers belong to a laboratory rig rather than a live plant; on the Co-Scientist side one architecture also failed physician evaluation. Better sensor data raised PLCBench success to 64.0 percent, while Co-Scientist reliability modules cut fabrication to 4 percent. PLCBench also records 98 episodes failing before basic reads succeeded; Co-Scientist reports obtaining thin films on a first attempt in materials science. PLCBench ran 240 episodes across five model families on real controllers. For readers the concrete development is industrial process risk being measured and lab automation being extended in the same week, with figures on each side coming from the authors' or the team's own measurement rather than independent replication, peer review, field deployment, outside audit or third-party confirmation.[1], [2]

References

  1. News sourcearXivAutonomous agents held a process off target in 75 of 240 runs on real controllers↩1↩2
  2. News sourceTHE DECODERCo-Scientist now plans the experiment, runs the lab equipment and writes the paper↩1↩2