The AI arena is free today

Open Superagent

AutomationBench v1.0.6

Implementation

Progress Over Time

Interactive timeline showing model performance evolution on AutomationBench v1.0.6

State-of-the-art frontier
Open
Proprietary

AutomationBench v1.0.6 Leaderboard

5 models
ContextCostLicense
11.0T1.0M$0.43 / $0.87
2309B1.0M$0.14 / $0.28
3
OpenAI
OpenAI
—1.1M$2.00 / $10.00
4
OpenAI
OpenAI
—1.1M$2.00 / $10.00
5—1.1M$0.10 / $0.50
Notice missing or incorrect data?
About this benchmark

What is AutomationBench v1.0.6?

AutomationBench v1.0.6 evaluates agent performance on multi-step workflow automation tasks.

AutomationBench v1.0.6 is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.5.

Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.

Current leaders

MiMo-V2.6-Pro from Xiaomi currently leads the AutomationBench v1.0.6 leaderboard with a score of 0.531 across 5 evaluated AI models.

1MiMo-V2.6-ProXiaomi53.1%
2MiMo-V2.6-FlashXiaomi52.3%
3GPT-6.1 SolOpenAI36.1%

FAQ

Common questions about the AutomationBench v1.0.6 benchmark and leaderboard.

What is the AutomationBench v1.0.6 benchmark?

AutomationBench v1.0.6 evaluates agent performance on multi-step workflow automation tasks.

What is the AutomationBench v1.0.6 leaderboard?

The AutomationBench v1.0.6 leaderboard ranks 5 AI models based on their performance on this benchmark. Currently, MiMo-V2.6-Pro by Xiaomi leads with a score of 0.531. The average score across all models is 0.391.

What is the highest AutomationBench v1.0.6 score?

The highest AutomationBench v1.0.6 score is 0.531, achieved by MiMo-V2.6-Pro from Xiaomi.

How many models are evaluated on AutomationBench v1.0.6?

5 models have been evaluated on the AutomationBench v1.0.6 benchmark, with 0 verified results and 5 self-reported results.

Where can I find the AutomationBench v1.0.6 dataset?

The AutomationBench v1.0.6 dataset is available at https://mimo.xiaomi.com/mimo-v2-6.

What categories does AutomationBench v1.0.6 cover?

AutomationBench v1.0.6 is categorized under reasoning, agents, and tool calling. The benchmark evaluates text models.

What's the difference between AutomationBench v1.0.6 and AutomationBench?

AutomationBench v1.0.6 is a variant of AutomationBench. See the AutomationBench leaderboard for the broader benchmark and per-model comparison.

What is the best open-source model on AutomationBench v1.0.6?

MiMo-V2.6-Pro by Xiaomi is the top-ranked open-source model on AutomationBench v1.0.6, with a score of 0.531 (rank #1).

Which model offers the best value on AutomationBench v1.0.6?

Among models scoring within 10% of the leader, MiMo-V2.6-Flash from Xiaomi is the cheapest, at $0.14 per million input tokens with a score of 0.523.

How recent are the AutomationBench v1.0.6 leaderboard results?

The AutomationBench v1.0.6 leaderboard was last updated in September 2026 and currently includes 5 evaluated models.