TAU3-Bench
Progress Over Time
Interactive timeline showing model performance evolution on TAU3-Bench
TAU3-Bench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Xiaomi | 1.0T | 1.0M | $0.43 / $0.87 | ||
| 2 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 3 | Zhipu AI | 754B | 200K | $1.40 / $4.40 | ||
| 4 | Alibaba Cloud / Qwen Team | 35B | — | — | ||
| 5 | 550B | — | — |
What is TAU3-Bench?
TAU3-Bench is a benchmark for evaluating general-purpose agent capabilities, testing models on multi-turn interactions with simulated user models, retrieval, and complex decision-making scenarios.
TAU3-Bench is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.7.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.
Current leaders
MiMo-V2.5-Pro from Xiaomi currently leads the TAU3-Bench leaderboard with a score of 0.729 across 5 evaluated AI models.
FAQ
Common questions about the TAU3-Bench benchmark and leaderboard.