Android Bench 2.0 Finds AI Coding Agents Still Struggle With Week-Long Tasks
Google’s Android Bench 2.0 shifts AI coding evaluation from simple patches to multi-day engineering work, where the top public pass rate falls to about 28 percent.

DeveloperTech reported that Google’s Android Bench 2.0 is testing frontier AI models on Android engineering work that can stretch over several days, making coding agents look far less complete than short bug-fix scoreboards imply.
The benchmark follows the Harbor framework and replaces narrow patch exercises with broader project work.
Models may have to refresh dependencies, create an app, deliver linked features or move software from a cross-platform base into native Android.
On earlier incremental-code tests, leading systems often reached about 91 percent.
The published benchmark table lists a best public pass rate of roughly 28 percent in the new long-horizon setting, with OpenAI’s GPT-6 Astra leading.
That drop is partly a measurement story.
Matthew McCullough, VP of Product Management for Android Developer, argued that a simple binary grade misses important progress on work that takes days.
The revised scoring can credit partial engineering success rather than erasing a run because one edge-case assertion failed.
A model might still receive meaningful credit after moving 40 screens to Jetpack Compose, organizing database tables and meeting 90 percent of the required behavior, even if it falls short of a strict pass.
Android Bench 2.0 evaluates completion through three lenses: whether the application works correctly, whether the interface matches the expected output and whether the change avoids regressions.
Deductions cover instruction violations and broken project constraints.
The model cards then place completion rates next to pass percentages and average computed cost per assignment, turning the benchmark into a closer read on usable work rather than a single win-loss count.
The results distinguish greenfield coding from maintenance inside an existing code base.
Fresh-file generation remains easier for the tested systems.
Refactoring demands more awareness of project structure, dependency relationships and architecture already in place.
The models are steadier on deterministic changes, including Java-to-Kotlin conversion, Retrofit-to-Ktor migration and ViewModel setup across projects larger than 125 files and 8,000 code lines.
The weak points show up when runtime behavior becomes less predictable.
Dependency-injection graphs that are not mapped cleanly, shifting framework versions and unreleased libraries all create reliability problems.
Cross-platform migration is another hard case: no model completed those tasks perfectly, and the strongest systems reached an 80 percent completion score rather than a full pass.
The benchmark also measures the agent environment around the model.
Early runs connected OpenAI’s GPT 5.6 Sol to Codex and Gemini 3.8 Flash to Google Antigravity.
DeveloperTech reported that prompt caching and tighter tool-window design reduced token use in complex sessions.
The leaderboard currently lists Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI’s GPT-6, Anthropic’s Fable 5.1, Kimi K3 and Qwen 3.8-Max, while later cycles are expected to test agent-and-model pairings across providers.




















