Eigen RadarAI
Analysis

WorldSolver tests coding agents against physical laws

The WorldSolver preprint asks coding agents to create numerical solvers for 168 physical phenomena. Evaluation checks executable code, visual agreement with the target behavior and physical rules throughout the simulation. The single study compares seven models. Kimi K2.7 Code submitted solvers on 99 per cent of tasks but generated valid simulation output on 38 per cent, with scenes and boundary conditions held fixed.

Artificial Intelligence··Evening
A shallow water tank with ripples in a bright laboratory, with a researcher seated at a computer behind it.

Solver submission and valid simulation diverged

WorldSolver is a single preprint evaluating coding agents that generate numerical solvers for physical phenomena. In its comparison of seven models, Kimi K2.7 Code submitted solvers on 99 per cent of tasks but obtained valid simulation output on 38 per cent. A solver computes how a physical system behaves over time. The tasks require agents to choose and implement that numerical method rather than use a ready-made physics engine.[1]

Scenes and boundary conditions stay fixed across 168 tasks

The evaluation builds 168 tasks from phenomena in 61 classic computer-graphics papers. Tasks include fluids and rigid bodies as well as deformable and complex materials. Other domains involve rods, strands, cloth and shells, alongside interactions between physical processes. Scene geometry, physical parameters, external inputs and rendering settings remain fixed. Agents can change only the solver code. Solver generation receives 1,800 seconds and execution of the resulting program receives 600 seconds.[1]

Physical checks examine every time step

Visual evaluation inspects 20 selected video frames, while physical checks use state variables at every time step. The task-specific checks test mass conservation, momentum–impulse consistency or whether objects penetrate one another. Criteria are hidden from agents. GPT-5.6-Sol has an average combined score of 48.7 per cent. The corresponding average for Claude Opus 5 is 46.7 per cent. Limited frame sampling can miss continuous-time errors, and a finite set of physical checks cannot cover every error.[1]

References

  1. News sourcearXivWorldSolver tests coding agents against the laws of physical simulation↩1↩2↩3