Eigen RadarAI
Analysis

Android Bench 2 scores how much of a long development task an agent completes

Google has expanded Android Bench to evaluate AI agents on development work that can take an engineer days. The new system scores partial completion, functionality, visual fidelity and regressions alongside overall task success. Dependency upgrades, app creation and porting join the task set. Google’s results distinguish stronger new-code generation from more difficult refactoring and runtime checks.

Artificial Intelligence··Midday
A phone in a testing rig beside a connected development computer.

Long tasks cover app creation, upgrades and porting

Google has released Android Bench 2.0, its framework for evaluating AI models and software agents on Android development. The update adds tasks that Google says can take an engineer several days or a week: upgrading dependencies, adding features, building an app from scratch and moving a cross-platform app to Android. It also introduces evaluations using agents supplied by model providers.[1]

Partial completion now survives a failing edge case

The scoring system measures completion using functionality, visual fidelity and the avoidance of regressions. Departures from instructions or structural constraints incur penalties. This replaces an all-or-nothing outcome in which a single failing edge-case assertion could fail a task despite many completed requirements. The framework’s original version had concentrated on incremental changes to existing repositories, including Android permissions, navigation and connectivity practices.[1]

Refactoring and runtime checks remain harder than new code

Google’s results find stronger performance in writing new code than in refactoring existing code. Defined transformations, such as converting Java to Kotlin, fare better; runtime validation, breaking framework changes and unfamiliar libraries remain difficult. InfoQ’s snapshot gives Claude Opus 5.5 a 32 per cent pass rate on long-horizon tasks and GPT 6 Astra 28 per cent. The best porting completion rate is 80 per cent, a separate measure from passing every requirement of a long task.[1]

References

  1. News sourceInfoQAndroid Bench 2 measures long development tasks and partial completion↩1↩2↩3