Toolathlon
Progress Over Time
Interactive timeline showing model performance evolution on Toolathlon
Toolathlon Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Meta | — | 1.0M | $1.25 / $4.25 | ||
| 2 | Kimi K3New Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 3 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 4 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 5 | Google | — | 1.0M | $1.50 / $9.00 | ||
| 6 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 7 | OpenAI | — | 1.0M | $2.50 / $15.00 | ||
| 8 | Anthropic | — | 1.0M | $3.00 / $15.00 | ||
| 9 | OpenAI | — | 1.1M | $1.00 / $6.00 | ||
| 10 | OpenAI | — | 1.1M | $2.50 / $15.00 | ||
| 11 | DeepSeek | 1.6T | 1.0M | $1.60 / $3.20 | ||
| 12 | ByteDance | — | — | — | ||
| 13 | Moonshot AI | 1.0T | 262K | $0.75 / $3.50 | ||
| 14 | Google | — | 1.0M | $0.50 / $3.00 | ||
| 15 | ByteDance | — | — | — | ||
| 16 | Tencent | 295B | — | — | ||
| 17 | Zhipu AI | 753B | 1.0M | $0.95 / $3.00 | ||
| 18 | DeepSeek | 284B | 1.0M | $0.10 / $0.20 | ||
| 19 | MiniMax | — | 205K | $0.30 / $1.20 | ||
| 19 | OpenAI | — | 400K | $1.75 / $14.00 | ||
| 21 | MiniMax | 230B | 1.0M | $0.30 / $1.20 | ||
| 22 | OpenAI | — | 400K | $0.75 / $4.50 | ||
| 23 | Zhipu AI | 754B | 200K | $1.40 / $4.40 | ||
| 24 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 25 | Alibaba Cloud / Qwen Team | 397B | — | — | ||
| 26 | OpenAI | — | 400K | $0.20 / $1.25 | ||
| 27 | DeepSeek | 685B | — | — | ||
| 27 | DeepSeek | 685B | — | — | ||
| 27 | DeepSeek | 685B | — | — | ||
| 30 | Alibaba Cloud / Qwen Team | 35B | — | — |
What is Toolathlon?
Tool Decathlon is a comprehensive benchmark for evaluating AI agents' ability to use multiple tools across diverse task categories. It measures proficiency in tool selection, sequencing, and execution across ten different tool-use scenarios.
Toolathlon is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 30 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.8.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.
Current leaders
Muse Spark 1.1 from Meta currently leads the Toolathlon leaderboard with a score of 0.756 across 30 evaluated AI models.
FAQ
Common questions about the Toolathlon benchmark and leaderboard.