SendTech Times
News
SYSTEMS SHIFT:

Android Bench 2.0 Finds AI Coding Agents Still Struggle With Week-Long Tasks

Newsroom brief

Google’s Android Bench 2.0 shifts AI coding evaluation from simple patches to multi-day engineering work, where the top public pass rate falls to about 28 percent.

Verified against source materialEdited by SendTech Times AI & Enterprise DeskSource: DeveloperTech
Android Bench 2.0 Finds AI Coding Agents Still Struggle With Week-Long Tasks
Image source: DeveloperTech

DeveloperTech reported that Google’s Android Bench 2.0 is testing frontier AI models on Android engineering work that can stretch over several days, making coding agents look far less complete than short bug-fix scoreboards imply.

The benchmark follows the Harbor framework and replaces narrow patch exercises with broader project work.

Models may have to refresh dependencies, create an app, deliver linked features or move software from a cross-platform base into native Android.

On earlier incremental-code tests, leading systems often reached about 91 percent.

The published benchmark table lists a best public pass rate of roughly 28 percent in the new long-horizon setting, with OpenAI’s GPT-6 Astra leading.

That drop is partly a measurement story.

Matthew McCullough, VP of Product Management for Android Developer, argued that a simple binary grade misses important progress on work that takes days.

The revised scoring can credit partial engineering success rather than erasing a run because one edge-case assertion failed.

A model might still receive meaningful credit after moving 40 screens to Jetpack Compose, organizing database tables and meeting 90 percent of the required behavior, even if it falls short of a strict pass.

Android Bench 2.0 evaluates completion through three lenses: whether the application works correctly, whether the interface matches the expected output and whether the change avoids regressions.

Deductions cover instruction violations and broken project constraints.

The model cards then place completion rates next to pass percentages and average computed cost per assignment, turning the benchmark into a closer read on usable work rather than a single win-loss count.

The results distinguish greenfield coding from maintenance inside an existing code base.

Fresh-file generation remains easier for the tested systems.

Refactoring demands more awareness of project structure, dependency relationships and architecture already in place.

The models are steadier on deterministic changes, including Java-to-Kotlin conversion, Retrofit-to-Ktor migration and ViewModel setup across projects larger than 125 files and 8,000 code lines.

The weak points show up when runtime behavior becomes less predictable.

Dependency-injection graphs that are not mapped cleanly, shifting framework versions and unreleased libraries all create reliability problems.

Cross-platform migration is another hard case: no model completed those tasks perfectly, and the strongest systems reached an 80 percent completion score rather than a full pass.

The benchmark also measures the agent environment around the model.

Early runs connected OpenAI’s GPT 5.6 Sol to Codex and Gemini 3.8 Flash to Google Antigravity.

DeveloperTech reported that prompt caching and tighter tool-window design reduced token use in complex sessions.

The leaderboard currently lists Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI’s GPT-6, Anthropic’s Fable 5.1, Kimi K3 and Qwen 3.8-Max, while later cycles are expected to test agent-and-model pairings across providers.

Share this article
inXf

Related articles

More
Z.ai GLM-5.2 Pushes Open Coding Models Into Longer Workflows
AI

Z.ai GLM-5.2 Pushes Open Coding Models Into Longer Workflows

Z.ai released GLM-5.2 under an MIT license with a one million-token context window, coding-agent benchmarks and self-hosting options, putting long-context software engineering back into the open-model race.

GitHub And Google Back ARD As AI Agents Search For Tools
AI

GitHub And Google Back ARD As AI Agents Search For Tools

GitHub, Google, Microsoft and other companies are backing Agentic Resource Discovery, a specification meant to help AI agents find, verify and connect to tools, skills, MCP servers and other resources without hard-coded integrations.

Block’s Builderbot Shows Where AI Coding Tools Hit The Enterprise Wall
AI

Block’s Builderbot Shows Where AI Coding Tools Hit The Enterprise Wall

Block says its Builderbot framework coordinates AI agents across internal repositories, Slack threads, issue trackers and continuous-integration workflows. The company says the system runs over 200,000 commands each day, merges about 1,500 pull requests each week and accounts for roughly fifteen percent of company code changes. The stronger claim is not code generation alone. Block is testing whether agentic software work can handle permissions, context, CI failures and customer-data isolation inside a large engineering organisation.

AWS Workflow Lets GitHub Actions Block Failed AI Agent Checks
AI

AWS Workflow Lets GitHub Actions Block Failed AI Agent Checks

AWS has published a reference workflow for testing AI agents in GitHub Actions, using AgentCore Evaluations to score traces, block risky merges and expose runtime and judge-model tradeoffs.

Logitech Aims MX Keypad At AI Coding With 135 Programmable Shortcuts
AI

Logitech Aims MX Keypad At AI Coding With 135 Programmable Shortcuts

Logitech launched the MX Keypad for developers, combining nine LCD keys, 15 pages of shortcuts, GitHub Copilot integration and AI-agent controls in a Rs 15,999 USB-C device.

Gemini 3.7 Flash Puts Google Agent Pricing On Trial
AI

Gemini 3.7 Flash Puts Google Agent Pricing On Trial

Google is rolling out Gemini 3.7 Flash with temporary API rates, stronger coding and workflow benchmarks, and a 2027 return to full pricing that leaves enterprises to test cost per completed task.

Xiaomi MiMo Code Tests Long-Horizon AI Coding Inside the Terminal
AI

Xiaomi MiMo Code Tests Long-Horizon AI Coding Inside the Terminal

Xiaomi has open-sourced MiMo Code V0.1.0, a terminal-native AI programming assistant built for long agentic software workflows. Internal testing with 576 developers and tasks exceeding 200 steps positions the release as a direct challenge to existing coding agents such as Claude Code.

OpenAI To Cut Cursor Model Access After SpaceX Acquisition
AI

OpenAI To Cut Cursor Model Access After SpaceX Acquisition

OpenAI plans to end Cursor’s access to its models on Nov. 12, 2026 after SpaceX acquired the AI coding startup, narrowing one model option for developers.

Keep Reading

More Stories

Latest
Kepler Targets 2027 Production for HBM Replacement MemoryCloud & Data CentersOct 6, 2026Kepler Targets 2027 Production for HBM Replacement MemoryEE Times reports that Kepler Computing is preparing 3D ferroelectric memory for 2027 production, promising higher capacity and bandwidth per watt while limiting reliance on advanced-node lithography.Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceCapital & PolicyOct 6, 2026Yokogawa Opens Singapore Hub For Industrial Cyber ResilienceYokogawa Engineering Asia has launched a Singapore center focused on OT cyber resilience, training, response planning and recovery coordination for Southeast Asia, Oceania and Taiwan.ClickFix Attack Uses Browser Cache To Hide Malware PayloadCybersecurityOct 6, 2026ClickFix Attack Uses Browser Cache To Hide Malware PayloadMicrosoft Threat Intelligence traced a ClickFix cache-smuggling method that preloads malware into browser caches, then uses file size checks and a pasted Run command to launch later credential-theft stages.VOA Tests Six-Month Startup Buildout Before Funding DecisionsFintech & Digital PaymentsOct 6, 2026VOA Tests Six-Month Startup Buildout Before Funding DecisionsTechCabal’s interview with VOA Venture Partners founder Victoria Olayide Adesanya describes a six-month build programme that lets the firm work inside African financial-infrastructure startups before deciding whether to invest.Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCrypto/Web3Oct 6, 2026Bitcoin Holds $86,000 As Dollar Index Hits 18-Month HighCoinDesk reported that bitcoin stayed near $86,000 while the U.S. Dollar Index reached about 102.5, with U.S. rate expectations and European political risks strengthening the dollar backdrop.Google Freezes OSS Bug Bounty Reports After AI Submission FloodCybersecurityOct 6, 2026Google Freezes OSS Bug Bounty Reports After AI Submission FloodGoogle has stopped accepting new product vulnerability reports in its OSS VRP after invalid automated submissions swamped reviewers, while older reports and some Cloud VRP routes remain open.Fleuret AI Raises €4M For Continuous AI Pentesting PlatformCybersecurityOct 6, 2026Fleuret AI Raises €4M For Continuous AI Pentesting PlatformTech.eu reported that French startup Fleuret AI raised €4 million in pre-seed funding to develop an agentic-AI platform that turns penetration testing into a continuous security process.GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%Fintech & Digital PaymentsOct 6, 2026GFT Analysis Says AI Documentation Can Cut Maintenance Work 30%A GFT Technologies analysis says AI-linked software documentation can cut maintenance effort and speed developer onboarding when knowledge assets stay synchronized with code changes.Schneider Electric Lines Up $22.6 Billion PTC DealAIOct 5, 2026Schneider Electric Lines Up $22.6 Billion PTC DealSchneider Electric plans to buy PTC in a cash transaction valuing the US engineering software provider’s equity at about $22.6 billion, adding product-lifecycle software to its industrial AI push.Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueCapital & PolicyOct 5, 2026Aggarwal Pledges Ola Electric Stake To Fund ₹1,000 Cr Rights IssueOla Electric founder Bhavish Aggarwal pledged 20 Cr shares to finance his participation in a rights issue that forms part of a larger ₹1,500 Cr fundraising plan.Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseAIOct 5, 2026Natrona Schools AI Review Puts Student Privacy Ahead Of Classroom Tool UseNatrona County trustees questioned whether teacher AI tools expose student data, even as existing district rules already ban unauthorized generative AI use by students.AMD Prices 256-Core EPYC 9996 At $14,904 For Server BuyersChips & SemiconductorsOct 5, 2026AMD Prices 256-Core EPYC 9996 At $14,904 For Server BuyersTechRadar reports that AMD’s 6th Gen EPYC 9006 “Venice” lineup includes a 256-core EPYC 9996 with 512 threads, 1GB of L3 cache, a 600W default power rating and a $14,904 list price for 1,000-unit orders.