Toolathlon

Progress Over Time

Interactive timeline showing model performance evolution on Toolathlon

State-of-the-art frontier
Open
Proprietary

Toolathlon Leaderboard

30 models
ContextCostLicense
11.0M$1.25 / $4.25
2
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
31.0M$5.00 / $25.00
41.1M$5.00 / $30.00
51.0M$1.50 / $9.00
6
OpenAI
OpenAI
1.1M$5.00 / $30.00
7
OpenAI
OpenAI
1.0M$2.50 / $15.00
81.0M$3.00 / $15.00
91.1M$1.00 / $6.00
101.1M$2.50 / $15.00
111.6T1.0M$1.60 / $3.20
12
ByteDance
ByteDance
13
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
141.0M$0.50 / $3.00
15
16
Tencent
Tencent
295B
17
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
18284B1.0M$0.10 / $0.20
19205K$0.30 / $1.20
19
OpenAI
OpenAI
400K$1.75 / $14.00
21230B1.0M$0.30 / $1.20
22400K$0.75 / $4.50
23
Zhipu AI
Zhipu AI
754B200K$1.40 / $4.40
24
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
25
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B
26400K$0.20 / $1.25
27685B
27685B
27685B
30
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
Notice missing or incorrect data?
About this benchmark

What is Toolathlon?

Tool Decathlon is a comprehensive benchmark for evaluating AI agents' ability to use multiple tools across diverse task categories. It measures proficiency in tool selection, sequencing, and execution across ten different tool-use scenarios.

Toolathlon is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 30 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.8.

Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.

Current leaders

Muse Spark 1.1 from Meta currently leads the Toolathlon leaderboard with a score of 0.756 across 30 evaluated AI models.

1Muse Spark 1.1Meta75.6%
2Kimi K3Moonshot AI73.2%
3Claude Opus 4.8Anthropic59.9%

FAQ

Common questions about the Toolathlon benchmark and leaderboard.

What is the Toolathlon benchmark?

Tool Decathlon is a comprehensive benchmark for evaluating AI agents' ability to use multiple tools across diverse task categories. It measures proficiency in tool selection, sequencing, and execution across ten different tool-use scenarios.

What is the Toolathlon leaderboard?

The Toolathlon leaderboard ranks 30 AI models based on their performance on this benchmark. Currently, Muse Spark 1.1 by Meta leads with a score of 0.756. The average score across all models is 0.485.

What is the highest Toolathlon score?

The highest Toolathlon score is 0.756, achieved by Muse Spark 1.1 from Meta.

How many models are evaluated on Toolathlon?

30 models have been evaluated on the Toolathlon benchmark, with 0 verified results and 30 self-reported results.

What categories does Toolathlon cover?

Toolathlon is categorized under reasoning, agents, and tool calling. The benchmark evaluates text models.

What is the best open-source model on Toolathlon?

Kimi K3 by Moonshot AI is the top-ranked open-source model on Toolathlon, with a score of 0.732 (rank #2).

Which model offers the best value on Toolathlon?

Among models scoring within 10% of the leader, Muse Spark 1.1 from Meta is the cheapest, at $1.25 per million input tokens with a score of 0.756.

How recent are the Toolathlon leaderboard results?

The Toolathlon leaderboard was last updated in July 2026 and currently includes 30 evaluated models.