Toolathlon-Verified
Progress Over Time
Interactive timeline showing model performance evolution on Toolathlon-Verified
Toolathlon-Verified Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $4.00 / $20.00 | ||
| 1 | Anthropic | — | 1.0M | $2.00 / $10.00 | ||
| 3 | Xiaomi | 1.0T | 1.0M | $0.43 / $0.87 | ||
| 4 | Xiaomi | 309B | 1.0M | $0.14 / $0.28 |
What is Toolathlon-Verified?
Verified Toolathlon evaluation of agent tool-use capability, as reported in the MiMo-V2.6 release.
Toolathlon-Verified is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.
Current leaders
Claude Opus 5.5 from Anthropic currently leads the Toolathlon-Verified leaderboard with a score of 0.778 across 4 evaluated AI models.
FAQ
Common questions about the Toolathlon-Verified benchmark and leaderboard.