Tau3 Banking
Progress Over Time
Interactive timeline showing model performance evolution on Tau3 Banking
Tau3 Banking Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Shanghai AI Laboratory | 744B | — | — | ||
| 2 | InclusionAI | 124B | 262K | $0.06 / $0.18 | ||
| 3 | xAI | — | 500K | $2.00 / $6.00 | ||
| 4 | OpenAI | — | — | — | ||
| 5 | InclusionAI | 124B | 131K | $0.06 / $0.18 | ||
| 6 | Thinking Machines Lab | 975B | 524K | $0.95 / $4.05 | ||
| 7 | Meta | 30B | 131K | $0.30 / $1.20 | ||
| 8 | Upstage | — | 524K | $0.30 / $1.20 | ||
| 9 | Thinking Machines Lab | 276B | 524K | $0.30 / $1.20 | ||
| 10 | Mistral AI | 128B | 256K | $1.50 / $7.50 | ||
| 11 | 30B | 262K | $0.08 / $0.20 | |||
| 12 | Liquid AI | 3B | — | — |
What is Tau3 Banking?
τ³-Bench banking domain evaluates agentic models on multi-turn, tool-using customer-support scenarios in a simulated retail banking environment.
Tau3 Banking is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 12 models on this benchmark, scored on a 0–1 scale. The current average is 0.2, with the leader at 0.4.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.
Current leaders
Atria Dawn Preview from Shanghai AI Laboratory currently leads the Tau3 Banking leaderboard with a score of 0.412 across 12 evaluated AI models.
FAQ
Common questions about the Tau3 Banking benchmark and leaderboard.