The AI arena is free today

Open Superagent

Toolathlon

Progress Over Time

Interactive timeline showing model performance evolution on Toolathlon

State-of-the-art frontier
Open
Proprietary

Toolathlon Leaderboard

42 models
ContextCostLicense
1320B1.0M$0.15 / $0.50
21.0M$1.25 / $4.25
31.6T1.0M$0.43 / $0.87
3770B
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B1.0M$0.15 / $0.47
7
Moonshot AI
Moonshot AI
2.8T1.0M$2.85 / $14.25
8
Zhipu AI
Zhipu AI
753B1.0M$1.20 / $4.00
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
10304B1.0M$0.06 / $0.18
111.0M$5.00 / $25.00
121.1M$5.00 / $30.00
131.0M$1.50 / $9.00
14
OpenAI
OpenAI
1.1M$5.00 / $30.00
15
OpenAI
OpenAI
1.0M$2.50 / $15.00
16
Thinking Machines Lab
Thinking Machines Lab
276B524K$0.30 / $1.20
171.0M$2.00 / $10.00
181.1M$0.20 / $1.20
191.1M$2.00 / $12.00
201.6T1.0M$1.30 / $2.60
21
ByteDance
ByteDance
22
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
23
Poolside
Poolside
118B1.0M$0.10 / $0.20
241.0M$0.50 / $3.00
25
26
Tencent
Tencent
295B262K$0.14 / $0.58
27
Zhipu AI
Zhipu AI
753B1.0M$0.75 / $2.40
28284B1.0M$0.09 / $0.18
29
OpenAI
OpenAI
400K$1.75 / $14.00
29205K$0.30 / $1.20
31230B1.0M$0.30 / $1.20
31284B1.0M$0.09 / $0.18
33400K$0.75 / $4.50
34
Zhipu AI
Zhipu AI
754B203K$1.05 / $3.50
35
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
36
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B262K$0.45 / $3.00
37400K$0.20 / $1.25
38685B
38685B
38685B164K$0.26 / $0.38
41
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.10 / $0.95
42
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0T256K$1.20 / $6.00
Notice missing or incorrect data?
About this benchmark

What is Toolathlon?

Tool Decathlon is a comprehensive benchmark for evaluating AI agents' ability to use multiple tools across diverse task categories. It measures proficiency in tool selection, sequencing, and execution across ten different tool-use scenarios.

Toolathlon is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 42 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.8.

Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.

Current leaders

GLM-5.3-Flash from Zhipu AI currently leads the Toolathlon leaderboard with a score of 0.784 across 42 evaluated AI models.

1GLM-5.3-FlashZhipu AI78.4%
2Muse Spark 1.1Meta75.6%
3DeepSeek-V4-Pro-0813DeepSeek74.1%

FAQ

Common questions about the Toolathlon benchmark and leaderboard.

What is the Toolathlon benchmark?

Tool Decathlon is a comprehensive benchmark for evaluating AI agents' ability to use multiple tools across diverse task categories. It measures proficiency in tool selection, sequencing, and execution across ten different tool-use scenarios.

What is the Toolathlon leaderboard?

The Toolathlon leaderboard ranks 42 AI models based on their performance on this benchmark. Currently, GLM-5.3-Flash by Zhipu AI leads with a score of 0.784. The average score across all models is 0.526.

What is the highest Toolathlon score?

The highest Toolathlon score is 0.784, achieved by GLM-5.3-Flash from Zhipu AI.

How many models are evaluated on Toolathlon?

42 models have been evaluated on the Toolathlon benchmark, with 0 verified results and 42 self-reported results.

What categories does Toolathlon cover?

Toolathlon is categorized under reasoning, agents, and tool calling. The benchmark evaluates text models.

What is the best open-source model on Toolathlon?

GLM-5.3-Flash by Zhipu AI is the top-ranked open-source model on Toolathlon, with a score of 0.784 (rank #1).

Which model offers the best value on Toolathlon?

Among models scoring within 10% of the leader, GLM-5.3-Flash from Zhipu AI is the cheapest, at $0.15 per million input tokens with a score of 0.784.

How recent are the Toolathlon leaderboard results?

The Toolathlon leaderboard was last updated in September 2026 and currently includes 42 evaluated models.