The AI arena is free today

Open Superagent

Toolathlon-Verified

Implementation

Progress Over Time

Interactive timeline showing model performance evolution on Toolathlon-Verified

State-of-the-art frontier
Open
Proprietary

Toolathlon-Verified Leaderboard

4 models
ContextCostLicense
1—1.0M$4.00 / $20.00
1—1.0M$2.00 / $10.00
31.0T1.0M$0.43 / $0.87
4309B1.0M$0.14 / $0.28
Notice missing or incorrect data?
About this benchmark

What is Toolathlon-Verified?

Verified Toolathlon evaluation of agent tool-use capability, as reported in the MiMo-V2.6 release.

Toolathlon-Verified is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.

Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.

Current leaders

Claude Opus 5.5 from Anthropic currently leads the Toolathlon-Verified leaderboard with a score of 0.778 across 4 evaluated AI models.

1Claude Opus 5.5Anthropic77.8%
1Claude Sonnet 5.5Anthropic77.8%
3MiMo-V2.6-ProXiaomi76.9%

FAQ

Common questions about the Toolathlon-Verified benchmark and leaderboard.

What is the Toolathlon-Verified benchmark?

Verified Toolathlon evaluation of agent tool-use capability, as reported in the MiMo-V2.6 release.

What is the Toolathlon-Verified leaderboard?

The Toolathlon-Verified leaderboard ranks 4 AI models based on their performance on this benchmark. Currently, Claude Opus 5.5 by Anthropic leads with a score of 0.778. The average score across all models is 0.765.

What is the highest Toolathlon-Verified score?

The highest Toolathlon-Verified score is 0.778, achieved by Claude Opus 5.5 from Anthropic.

How many models are evaluated on Toolathlon-Verified?

4 models have been evaluated on the Toolathlon-Verified benchmark, with 0 verified results and 4 self-reported results.

Where can I find the Toolathlon-Verified dataset?

The Toolathlon-Verified dataset is available at https://mimo.xiaomi.com/mimo-v2-6.

What categories does Toolathlon-Verified cover?

Toolathlon-Verified is categorized under reasoning, agents, and tool calling. The benchmark evaluates text models.

What's the difference between Toolathlon-Verified and Toolathlon?

Toolathlon-Verified is a variant of Toolathlon. See the Toolathlon leaderboard for the broader benchmark and per-model comparison.

What is the best open-source model on Toolathlon-Verified?

MiMo-V2.6-Pro by Xiaomi is the top-ranked open-source model on Toolathlon-Verified, with a score of 0.769 (rank #3).

Which model offers the best value on Toolathlon-Verified?

Among models scoring within 10% of the leader, MiMo-V2.6-Flash from Xiaomi is the cheapest, at $0.14 per million input tokens with a score of 0.736.

How recent are the Toolathlon-Verified leaderboard results?

The Toolathlon-Verified leaderboard was last updated in October 2026 and currently includes 4 evaluated models.