AMD agents clear most driver issues while a research redo fails author review
AMD says agents now close more than 75 per cent of a driver issue queue. A Princeton redo of two unpublished NeurIPS papers failed author review. Claude Code added an unmeasured design step before the code.
Artificial Intelligence··Evening
AMD reports agents now close most of a driver queue
Writing in IEEE Spectrum on 17 August, AMD senior vice president Andrej Zdravkovic reported that agents fixing issues in the company's Radeon Software eXperience resolved 6 per cent of them when the work started in October 2025 and more than 75 per cent by June 2026. He attributes the gain to refining the objectives given to the agents instead of retraining the underlying models. AMD's one objective metric is the share of source code generated by AI, counted only after it passes review and testing: past 20 per cent at the start of this year, heading toward 50 per cent across the codebase, and above 80 per cent in some components. The company reports a 30 per cent overall productivity gain, ahead of the 25 per cent it had hoped for over two or three years. These are AMD's own measurements, published in an opinion piece by one of its executives, with no independent audit and no figures on review time, defect rates or rollback burden.[1]
A six-day research redo fails both original authors
A study led by Peter Kirgis and Sayash Kapoor at Princeton gave Claude Opus 4.8 six days, 3,000 dollars in interface credits and GPU access to redo the research behind two unpublished NeurIPS 2026 papers. The original authors rejected both attempts. Kapoor's reading is that the agents were capable of all the engineering the research required and unambiguously bad at the research itself: they ran poorly designed experiments and abandoned ambitious hypotheses on very thin data. The result speaks to recursive self-improvement, the idea that AI systems can speed up their own development. Two caveats sit on it: only two papers were evaluated, and the graders knew they were assessing AI-generated work, which can bias the judgement. Kapoor frames the remaining question as whether creative leaps are strictly required for that loop to start.[2]
Claude Code drafts interface mockups before the code
Anthropic released an early preview of a /design command for Claude Code on 18 August. Typing '/design a few options for {feature}' returns several drafts laid out as artboards; the developer picks one, edits it and then builds it out in the same session. The command reads the existing codebase, matches the project's current interface style and produces shareable mockups as Artifacts. The chosen design carries over into the build step, though designs have to be saved by hand for now. Access comes through the claude update command in the terminal or the desktop app. This is an early preview rather than a generally available product, and Anthropic has published no measurement of how the step changes review or rework.[3]
Related columns
For more information on this topic, you can read the related columns.