AutomationBench v1.0.6
Progress Over Time
Interactive timeline showing model performance evolution on AutomationBench v1.0.6
AutomationBench v1.0.6 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Xiaomi | 1.0T | 1.0M | $0.43 / $0.87 | ||
| 2 | Xiaomi | 309B | 1.0M | $0.14 / $0.28 | ||
| 3 | GPT-6.1 SolNew OpenAI | — | 1.1M | $2.00 / $10.00 | ||
| 4 | OpenAI | — | 1.1M | $2.00 / $10.00 | ||
| 5 | OpenAI | — | 1.1M | $0.10 / $0.50 |
What is AutomationBench v1.0.6?
AutomationBench v1.0.6 evaluates agent performance on multi-step workflow automation tasks.
AutomationBench v1.0.6 is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.5.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.
Current leaders
MiMo-V2.6-Pro from Xiaomi currently leads the AutomationBench v1.0.6 leaderboard with a score of 0.531 across 5 evaluated AI models.
FAQ
Common questions about the AutomationBench v1.0.6 benchmark and leaderboard.