AutomationBench
Progress Over Time
Interactive timeline showing model performance evolution on AutomationBench
AutomationBench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Kimi K3New Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 2 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 3 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 4 | OpenAI | — | 1.1M | $2.50 / $15.00 | ||
| 5 | OpenAI | — | 1.1M | $1.00 / $6.00 | ||
| 6 | Anthropic | — | 1.0M | $3.00 / $15.00 |
Sub-benchmarks
What is AutomationBench?
AutomationBench is a tool-use benchmark that evaluates AI agents on automating real-world workflows, testing their ability to orchestrate tools and complete multi-step automation tasks.
AutomationBench is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 6 models on this benchmark, scored on a 0–1 scale. The current average is 0.2, with the leader at 0.3.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.
Current leaders
Kimi K3 from Moonshot AI currently leads the AutomationBench leaderboard with a score of 0.308 across 6 evaluated AI models.
FAQ
Common questions about the AutomationBench benchmark and leaderboard.