Terminal-Bench 4.0
Progress Over Time
Interactive timeline showing model performance evolution on Terminal-Bench 4.0
Terminal-Bench 4.0 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | — | — | ||
| 2 | GPT-6 AstraNew OpenAI | — | 1.1M | $10.00 / $50.00 | ||
| 3 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 4 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 5 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 6 | Zhipu AI | 753B | 1.0M | $1.20 / $4.00 | ||
| 7 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 8 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 9 | OpenAI | — | 1.1M | $2.00 / $12.00 | ||
| 10 | xAI | — | 500K | $2.00 / $6.00 | ||
| 11 | Google | — | 1.0M | $0.75 / $3.75 | ||
| 12 | OpenAI | — | 1.1M | $0.20 / $1.20 | ||
| 13 | Anthropic | — | 1.0M | $2.00 / $10.00 | ||
| 13 | xAI | — | 500K | $2.00 / $6.00 |
What is Terminal-Bench 4.0?
Terminal-Bench 4.0 evaluates AI agents on real-world work in terminal and command-line containerized environments across 66 tasks, with emphasis on science-adjacent and frontier engineering problems. Relative to earlier Terminal-Bench releases, 4.0 increases timeouts and adaptively raises RAM/CPU on selected tasks to reduce harness and resource confounds.
Terminal-Bench 4.0 is a text benchmark evaluating models on reasoning, agents, code, and tool calling tasks. LLM Stats tracks 14 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.6.
Compare leaders on the best AI for reasoning, best AI for agents, best AI for code and best AI for tool calling leaderboards.
Current leaders
Claude Mythos 5.1 from Anthropic currently leads the Terminal-Bench 4.0 leaderboard with a score of 0.609 across 14 evaluated AI models.
FAQ
Common questions about the Terminal-Bench 4.0 benchmark and leaderboard.