# Google launches Android Bench 2.0 with long‑horizon tasks and new scoring

> The updated benchmark adds multi‑day development tasks, agent‑based evaluation and a graded scoring system to measure AI assistants for Android coding.

Oossa · 2026-10-09 · https://oossa.com/en/google-launches-android-bench-2-0-with-long-horizon-tasks-and-new-scoring

Google has released Android Bench 2.0, an upgrade to its Android‑development benchmark for AI models. The new version adds long‑horizon tasks that can take an engineer days or a week, and it evaluates agents rather than just code snippets. Scoring is now continuous, judging functionality, visual fidelity and instruction adherence instead of a simple pass/fail. The dashboard shows Claude Opus 5.5 leading with a 32% pass rate on the hardest tasks.

## The facts

- Android Bench 2.0 announced on 2026-10-09 ("Fri Oct 09 2026 19:00:00 GMT+0200").
- Claude Opus 5.5 topped the LHT leaderboard with a 32% pass rate.

## Why it matters

Developers can use the benchmark to see which AI assistants are most reliable for complex Android projects.

## Sources & references

1. [Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring](https://www.infoq.com/news/2026/10/android-bench-2/?utm_campaign=infoq_content&utm_source=infoq&utm_medium=feed&utm_term=AI%2C+ML+%26+Data+Engineering) – InfoQ, 2026-10-09

Last updated: 2026-10-09
