The AI arena is free today

Open Superagent

TAU3-Bench

Progress Over Time

Interactive timeline showing model performance evolution on TAU3-Bench

State-of-the-art frontier
Open
Proprietary

TAU3-Bench Leaderboard

5 models
ContextCostLicense
11.0T1.0M$0.43 / $0.87
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
3
Zhipu AI
Zhipu AI
754B200K$1.40 / $4.40
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
5550B
Notice missing or incorrect data?
About this benchmark

What is TAU3-Bench?

TAU3-Bench is a benchmark for evaluating general-purpose agent capabilities, testing models on multi-turn interactions with simulated user models, retrieval, and complex decision-making scenarios.

TAU3-Bench is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.7.

Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.

Current leaders

MiMo-V2.5-Pro from Xiaomi currently leads the TAU3-Bench leaderboard with a score of 0.729 across 5 evaluated AI models.

1MiMo-V2.5-ProXiaomi72.9%
2Qwen3.6 PlusAlibaba Cloud / Qwen Team70.7%
3GLM-5.1Zhipu AI70.6%

FAQ

Common questions about the TAU3-Bench benchmark and leaderboard.

What is the TAU3-Bench benchmark?

TAU3-Bench is a benchmark for evaluating general-purpose agent capabilities, testing models on multi-turn interactions with simulated user models, retrieval, and complex decision-making scenarios.

What is the TAU3-Bench leaderboard?

The TAU3-Bench leaderboard ranks 5 AI models based on their performance on this benchmark. Currently, MiMo-V2.5-Pro by Xiaomi leads with a score of 0.729. The average score across all models is 0.608.

What is the highest TAU3-Bench score?

The highest TAU3-Bench score is 0.729, achieved by MiMo-V2.5-Pro from Xiaomi.

How many models are evaluated on TAU3-Bench?

5 models have been evaluated on the TAU3-Bench benchmark, with 0 verified results and 5 self-reported results.

What categories does TAU3-Bench cover?

TAU3-Bench is categorized under reasoning, agents, and tool calling. The benchmark evaluates text models.

What is the best open-source model on TAU3-Bench?

MiMo-V2.5-Pro by Xiaomi is the top-ranked open-source model on TAU3-Bench, with a score of 0.729 (rank #1).

Which model offers the best value on TAU3-Bench?

Among models scoring within 10% of the leader, MiMo-V2.5-Pro from Xiaomi is the cheapest, at $0.43 per million input tokens with a score of 0.729.

How recent are the TAU3-Bench leaderboard results?

The TAU3-Bench leaderboard was last updated in August 2026 and currently includes 5 evaluated models.
TAU3-Bench Leaderboard