Terminal-Bench 3.0
Progress Over Time
Interactive timeline showing model performance evolution on Terminal-Bench 3.0
Terminal-Bench 3.0 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Zhipu AI | 753B | 1.0M | $1.40 / $4.40 | ||
| 2 | xAI | — | 500K | $2.00 / $6.00 | ||
| 3 | Google | — | 1.0M | $0.75 / $3.75 |
What is Terminal-Bench 3.0?
Terminal-Bench 3.0 is a release of the Terminal-Bench benchmark that tests AI agents' ability to operate a computer via the terminal on real-world, end-to-end tasks.
Terminal-Bench 3.0 is a text benchmark evaluating models on reasoning, agents, code, and tool calling tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.2, with the leader at 0.3.
Compare leaders on the best AI for reasoning, best AI for agents, best AI for code and best AI for tool calling leaderboards.
Current leaders
GLM-5.3 from Zhipu AI currently leads the Terminal-Bench 3.0 leaderboard with a score of 0.283 across 3 evaluated AI models.
FAQ
Common questions about the Terminal-Bench 3.0 benchmark and leaderboard.