Google has released Android Bench 2.0, an upgrade to its Android‑development benchmark for AI models. The new version adds long‑horizon tasks that can take an engineer days or a week, and it evaluates agents rather than just code snippets. Scoring is now continuous, judging functionality, visual fidelity and instruction adherence instead of a simple pass/fail. The dashboard shows Claude Opus 5.5 leading with a 32% pass rate on the hardest tasks.
Why it matters
Developers can use the benchmark to see which AI assistants are most reliable for complex Android projects.