The AI arena is free today

Open Superagent

AndroidBench

Progress Over Time

Interactive timeline showing model performance evolution on AndroidBench

State-of-the-art frontier
Open
Proprietary

AndroidBench Leaderboard

1 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
Notice missing or incorrect data?
About this benchmark

What is AndroidBench?

AndroidBench evaluates coding agents on Android application development tasks.

AndroidBench is a multimodal benchmark evaluating models on agents, code, and tool calling tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.

Compare leaders on the best AI for agents, best AI for code and best AI for tool calling leaderboards.

Current leaders

Qwen3.8 Max from Alibaba Cloud / Qwen Team currently leads the AndroidBench leaderboard with a score of 0.751 across 1 evaluated AI models.

1Qwen3.8 MaxAlibaba Cloud / Qwen Team75.1%

FAQ

Common questions about the AndroidBench benchmark and leaderboard.

What is the AndroidBench benchmark?

AndroidBench evaluates coding agents on Android application development tasks.

What is the AndroidBench leaderboard?

The AndroidBench leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Qwen3.8 Max by Alibaba Cloud / Qwen Team leads with a score of 0.751. The average score across all models is 0.751.

What is the highest AndroidBench score?

The highest AndroidBench score is 0.751, achieved by Qwen3.8 Max from Alibaba Cloud / Qwen Team.

How many models are evaluated on AndroidBench?

1 models have been evaluated on the AndroidBench benchmark, with 0 verified results and 1 self-reported results.

What categories does AndroidBench cover?

AndroidBench is categorized under agents, code, and tool calling. The benchmark evaluates multimodal models.

What is the best open-source model on AndroidBench?

Qwen3.8 Max by Alibaba Cloud / Qwen Team is the top-ranked open-source model on AndroidBench, with a score of 0.751 (rank #1).

Which model offers the best value on AndroidBench?

Among models scoring within 10% of the leader, Qwen3.8 Max from Alibaba Cloud / Qwen Team is the cheapest, at $1.65 per million input tokens with a score of 0.751.

How recent are the AndroidBench leaderboard results?

The AndroidBench leaderboard was last updated in August 2026 and currently includes 1 evaluated models.