Terminal-Bench
Progress Over Time
Interactive timeline showing model performance evolution on Terminal-Bench
Terminal-Bench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 200K | $3.00 / $15.00 | ||
| 2 | MiniMax | 230B | 1.0M | $0.30 / $1.20 | ||
| 3 | Moonshot AI | 1.0T | — | — | ||
| 4 | MiniMax | 230B | 1.0M | $0.30 / $1.20 | ||
| 5 | Anthropic | — | — | — | ||
| 6 | Amazon | — | — | — | ||
| 7 | Anthropic | — | 200K | $1.00 / $5.00 | ||
| 8 | Zhipu AI | 357B | 203K | $0.50 / $2.00 | ||
| 9 | Meituan | 560B | — | — | ||
| 10 | Anthropic | — | — | — | ||
| 11 | DeepSeek | 685B | — | — | ||
| 12 | Zhipu AI | 355B | — | — | ||
| 13 | Anthropic | — | — | — | ||
| 14 | Anthropic | — | — | — | ||
| 15 | Meituan | 69B | 256K | $0.10 / $0.40 | ||
| 16 | Zhipu AI | 358B | 203K | $0.40 / $1.75 | ||
| 17 | Amazon | — | 1.0M | $0.30 / $2.50 | ||
| 18 | DeepSeek | 671B | 164K | $0.25 / $0.95 | ||
| 19 | Xiaomi | 309B | — | — | ||
| 20 | Zhipu AI | 106B | — | — | ||
| 20 | Moonshot AI | 1.0T | — | — | ||
| 22 | 120B | 262K | $0.09 / $0.40 | |||
| 23 | Moonshot AI | 1.0T | — | — | ||
| 24 | 32B | 262K | $0.05 / $0.20 | |||
| 25 | DeepSeek | 671B | 164K | $0.50 / $2.15 |
Sub-benchmarks
Terminal-Bench 2.0
Terminal-Bench 2.0 is an updated benchmark for testing AI agents' tool use ability to operate a computer via terminal. It evaluates how well models can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities.
Terminal-Bench 2.1
Terminal-Bench 2.1 is an updated release of the Terminal-Bench benchmark that tests AI agents' ability to operate a computer via the terminal. It evaluates how well models handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, data science workflows, and security tasks.
Terminal-Bench 3.0
Terminal-Bench 3.0 is a release of the Terminal-Bench benchmark that tests AI agents' ability to operate a computer via the terminal on real-world, end-to-end tasks.
Terminal-Bench 4.0
Terminal-Bench 4.0 evaluates AI agents on real-world work in terminal and command-line containerized environments across 66 tasks, with emphasis on science-adjacent and frontier engineering problems. Relative to earlier Terminal-Bench releases, 4.0 increases timeouts and adaptively raises RAM/CPU on selected tasks to reduce harness and resource confounds.
Terminal-Bench Hard
Terminal-Bench Hard is a harder terminal-agent benchmark variant evaluated with the Terminus-2 harness in Cohere's Command A+ and North Mini Code releases.
Terminal-Bench-Science 0.1
Terminal-Bench-Science 0.1 is a Stanford-led community benchmark of 70 tasks drawn from scientific research workflows across the life, physical, earth, mathematical, and engineering sciences. Tasks are authored and reviewed by scientists and researchers; scores are typically reported as accuracy with relatively large standard error due to strongly bimodal per-task outcomes.
What is Terminal-Bench?
Terminal-Bench is a benchmark for testing AI agents in real terminal environments. It evaluates how well agents can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities. The benchmark consists of a dataset of ~100 hand-crafted, human-verified tasks and an execution harness that connects language models to a terminal sandbox.
Terminal-Bench is a text benchmark evaluating models on reasoning, agents, and code tasks. LLM Stats tracks 25 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.5.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for code leaderboards.
Current leaders
Claude Sonnet 4.5 from Anthropic currently leads the Terminal-Bench leaderboard with a score of 0.500 across 25 evaluated AI models.
FAQ
Common questions about the Terminal-Bench benchmark and leaderboard.