AutomationBench
Progress Over Time
Interactive timeline showing model performance evolution on AutomationBench
AutomationBench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | DeepSeek | 763B | 1.0M | $0.22 / $0.66 | ||
| 2 | Shanghai AI Laboratory | 744B | — | — | ||
| 3 | Meta | — | 1.0M | $0.10 / $0.20 | ||
| 4 | Zhipu AI | 320B | 1.0M | $0.15 / $0.50 | ||
| 5 | Zhipu AI | 753B | 1.0M | $1.20 / $4.00 | ||
| 6 | OpenAI | — | 1.1M | $10.00 / $50.00 | ||
| 7 | Tencent | 770B | — | — | ||
| 8 | DeepSeek | 1.6T | 1.0M | $1.30 / $2.60 | ||
| 9 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 10 | Moonshot AI | 2.8T | 1.0M | $2.85 / $14.25 | ||
| 11 | Google | — | 1.0M | $0.75 / $3.75 | ||
| 12 | Alibaba Cloud / Qwen Team | 2.4T | 1.0M | $1.65 / $4.95 | ||
| 13 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 14 | DeepSeek | — | 1.0M | $0.44 / $1.32 | ||
| 15 | DeepSeek | 304B | 1.0M | $0.06 / $0.18 | ||
| 16 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 17 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 18 | OpenAI | — | 1.1M | $2.00 / $12.00 | ||
| 19 | OpenAI | — | 1.1M | $0.20 / $1.20 | ||
| 20 | Anthropic | — | 1.0M | $2.00 / $10.00 |
Sub-benchmarks
What is AutomationBench?
AutomationBench is a tool-use benchmark that evaluates AI agents on automating real-world workflows, testing their ability to orchestrate tools and complete multi-step automation tasks.
AutomationBench is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 20 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.5.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.
Current leaders
DeepSeek-V4.1-Flash from DeepSeek currently leads the AutomationBench leaderboard with a score of 0.548 across 20 evaluated AI models.
FAQ
Common questions about the AutomationBench benchmark and leaderboard.