Toolathlon
Progress Over Time
Interactive timeline showing model performance evolution on Toolathlon
Toolathlon Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Zhipu AI | 320B | 1.0M | $0.15 / $0.50 | ||
| 2 | Meta | — | 1.0M | $1.25 / $4.25 | ||
| 3 | DeepSeek | 1.6T | 1.0M | $0.43 / $0.87 | ||
| 3 | Tencent | 770B | — | — | ||
| 5 | Alibaba Cloud / Qwen Team | 125B | — | — | ||
| 5 | Alibaba Cloud / Qwen Team | 125B | 1.0M | $0.15 / $0.47 | ||
| 7 | Moonshot AI | 2.8T | 1.0M | $2.85 / $14.25 | ||
| 8 | Zhipu AI | 753B | 1.0M | $1.20 / $4.00 | ||
| 9 | Alibaba Cloud / Qwen Team | 2.4T | 1.0M | $1.65 / $4.95 | ||
| 10 | DeepSeek | 304B | 1.0M | $0.06 / $0.18 | ||
| 11 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 12 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 13 | Google | — | 1.0M | $1.50 / $9.00 | ||
| 14 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 15 | OpenAI | — | 1.0M | $2.50 / $15.00 | ||
| 16 | Thinking Machines Lab | 276B | 524K | $0.30 / $1.20 | ||
| 17 | Anthropic | — | 1.0M | $2.00 / $10.00 | ||
| 18 | OpenAI | — | 1.1M | $0.20 / $1.20 | ||
| 19 | OpenAI | — | 1.1M | $2.00 / $12.00 | ||
| 20 | DeepSeek | 1.6T | 1.0M | $1.30 / $2.60 | ||
| 21 | ByteDance | — | — | — | ||
| 22 | Moonshot AI | 1.0T | 262K | $0.75 / $3.50 | ||
| 23 | Poolside | 118B | 1.0M | $0.10 / $0.20 | ||
| 24 | Google | — | 1.0M | $0.50 / $3.00 | ||
| 25 | ByteDance | — | — | — | ||
| 26 | Tencent | 295B | 262K | $0.14 / $0.58 | ||
| 27 | Zhipu AI | 753B | 1.0M | $0.75 / $2.40 | ||
| 28 | DeepSeek | 284B | 1.0M | $0.09 / $0.18 | ||
| 29 | OpenAI | — | 400K | $1.75 / $14.00 | ||
| 29 | MiniMax | — | 205K | $0.30 / $1.20 | ||
| 31 | MiniMax | 230B | 1.0M | $0.30 / $1.20 | ||
| 31 | DeepSeek | 284B | 1.0M | $0.09 / $0.18 | ||
| 33 | OpenAI | — | 400K | $0.75 / $4.50 | ||
| 34 | Zhipu AI | 754B | 203K | $1.05 / $3.50 | ||
| 35 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 36 | Alibaba Cloud / Qwen Team | 397B | 262K | $0.45 / $3.00 | ||
| 37 | OpenAI | — | 400K | $0.20 / $1.25 | ||
| 38 | DeepSeek | 685B | — | — | ||
| 38 | DeepSeek | 685B | — | — | ||
| 38 | DeepSeek | 685B | 164K | $0.26 / $0.38 | ||
| 41 | Alibaba Cloud / Qwen Team | 35B | 262K | $0.10 / $0.95 | ||
| 42 | Alibaba Cloud / Qwen Team | 1.0T | 256K | $1.20 / $6.00 |
What is Toolathlon?
Tool Decathlon is a comprehensive benchmark for evaluating AI agents' ability to use multiple tools across diverse task categories. It measures proficiency in tool selection, sequencing, and execution across ten different tool-use scenarios.
Toolathlon is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 42 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.8.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.
Current leaders
GLM-5.3-Flash from Zhipu AI currently leads the Toolathlon leaderboard with a score of 0.784 across 42 evaluated AI models.
FAQ
Common questions about the Toolathlon benchmark and leaderboard.