Terminal-Bench 2.1
Progress Over Time
Interactive timeline showing model performance evolution on Terminal-Bench 2.1
Terminal-Bench 2.1 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 2 | Kimi K3New Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 3 | OpenAI | — | 1.1M | $2.50 / $15.00 | ||
| 4 | OpenAI | — | 1.1M | $1.00 / $6.00 | ||
| 5 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 6 | xAI | — | 500K | $2.00 / $6.00 | ||
| 7 | Zhipu AI | 753B | 1.0M | $0.95 / $3.00 | ||
| 8 | Tencent | 295B | — | — | ||
| 9 | ByteDance | — | — | — | ||
| 10 | ByteDance | — | — | — | ||
| 11 | MiniMax | — | 1.0M | $0.30 / $1.20 | ||
| 12 | 550B | — | — |
What is Terminal-Bench 2.1?
Terminal-Bench 2.1 is an updated release of the Terminal-Bench benchmark that tests AI agents' ability to operate a computer via the terminal. It evaluates how well models handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, data science workflows, and security tasks.
Terminal-Bench 2.1 is a text benchmark evaluating models on reasoning, agents, code, and tool calling tasks. LLM Stats tracks 12 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.9.
Compare leaders on the best AI for reasoning, best AI for agents, best AI for code and best AI for tool calling leaderboards.
Current leaders
GPT-5.6 Sol from OpenAI currently leads the Terminal-Bench 2.1 leaderboard with a score of 0.888 across 12 evaluated AI models.
FAQ
Common questions about the Terminal-Bench 2.1 benchmark and leaderboard.