The AI arena is free today

Open Superagent

Terminal-Bench

Progress Over Time

Interactive timeline showing model performance evolution on Terminal-Bench

State-of-the-art frontier
Open
Proprietary

Terminal-Bench Leaderboard

25 models
ContextCostLicense
1—200K$3.00 / $15.00
2230B1.0M$0.30 / $1.20
31.0T——
4
MiniMax
MiniMax
230B1.0M$0.30 / $1.20
5———
6———
7—200K$1.00 / $5.00
8
Zhipu AI
Zhipu AI
357B203K$0.50 / $2.00
9560B——
10
Anthropic
Anthropic
———
11685B——
12
Zhipu AI
Zhipu AI
355B——
13———
14———
1569B256K$0.10 / $0.40
16
Zhipu AI
Zhipu AI
358B203K$0.40 / $1.75
17—1.0M$0.30 / $2.50
18671B164K$0.25 / $0.95
19309B——
20
Zhipu AI
Zhipu AI
106B——
20
Moonshot AI
Moonshot AI
1.0T——
22120B262K$0.09 / $0.40
231.0T——
2432B262K$0.05 / $0.20
25671B164K$0.50 / $2.15
Notice missing or incorrect data?

Sub-benchmarks

Terminal-Bench 2.0

Terminal-Bench 2.0 is an updated benchmark for testing AI agents' tool use ability to operate a computer via terminal. It evaluates how well models can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities.

text•Max 1

Terminal-Bench 2.1

Terminal-Bench 2.1 is an updated release of the Terminal-Bench benchmark that tests AI agents' ability to operate a computer via the terminal. It evaluates how well models handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, data science workflows, and security tasks.

text•Max 1

Terminal-Bench 3.0

Terminal-Bench 3.0 is a release of the Terminal-Bench benchmark that tests AI agents' ability to operate a computer via the terminal on real-world, end-to-end tasks.

text•Max 1

Terminal-Bench 4.0

Terminal-Bench 4.0 evaluates AI agents on real-world work in terminal and command-line containerized environments across 66 tasks, with emphasis on science-adjacent and frontier engineering problems. Relative to earlier Terminal-Bench releases, 4.0 increases timeouts and adaptively raises RAM/CPU on selected tasks to reduce harness and resource confounds.

text•Max 1

Terminal-Bench Hard

Terminal-Bench Hard is a harder terminal-agent benchmark variant evaluated with the Terminus-2 harness in Cohere's Command A+ and North Mini Code releases.

text•Max 1

Terminal-Bench-Science 0.1

Terminal-Bench-Science 0.1 is a Stanford-led community benchmark of 70 tasks drawn from scientific research workflows across the life, physical, earth, mathematical, and engineering sciences. Tasks are authored and reviewed by scientists and researchers; scores are typically reported as accuracy with relatively large standard error due to strongly bimodal per-task outcomes.

text•Max 1
About this benchmark

What is Terminal-Bench?

Terminal-Bench is a benchmark for testing AI agents in real terminal environments. It evaluates how well agents can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities. The benchmark consists of a dataset of ~100 hand-crafted, human-verified tasks and an execution harness that connects language models to a terminal sandbox.

Terminal-Bench is a text benchmark evaluating models on reasoning, agents, and code tasks. LLM Stats tracks 25 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.5.

Compare leaders on the best AI for reasoning, best AI for agents and best AI for code leaderboards.

Current leaders

Claude Sonnet 4.5 from Anthropic currently leads the Terminal-Bench leaderboard with a score of 0.500 across 25 evaluated AI models.

1Claude Sonnet 4.5Anthropic50.0%
2MiniMax M2.1MiniMax47.9%
3Kimi K2-Thinking-0905Moonshot AI47.1%

FAQ

Common questions about the Terminal-Bench benchmark and leaderboard.

What is the Terminal-Bench benchmark?

Terminal-Bench is a benchmark for testing AI agents in real terminal environments. It evaluates how well agents can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities. The benchmark consists of a dataset of ~100 hand-crafted, human-verified tasks and an execution harness that connects language models to a terminal sandbox.

What is the Terminal-Bench leaderboard?

The Terminal-Bench leaderboard ranks 25 AI models based on their performance on this benchmark. Currently, Claude Sonnet 4.5 by Anthropic leads with a score of 0.500. The average score across all models is 0.347.

What is the highest Terminal-Bench score?

The highest Terminal-Bench score is 0.500, achieved by Claude Sonnet 4.5 from Anthropic.

How many models are evaluated on Terminal-Bench?

25 models have been evaluated on the Terminal-Bench benchmark, with 0 verified results and 25 self-reported results.

What categories does Terminal-Bench cover?

Terminal-Bench is categorized under reasoning, agents, and code. The benchmark evaluates text models.

Are there variants of Terminal-Bench?

Yes. Terminal-Bench has 6 related variants: Terminal-Bench 2.0, Terminal-Bench 2.1, Terminal-Bench 3.0, Terminal-Bench 4.0.

What is the best open-source model on Terminal-Bench?

MiniMax M2.1 by MiniMax is the top-ranked open-source model on Terminal-Bench, with a score of 0.479 (rank #2).

Which model offers the best value on Terminal-Bench?

Among models scoring within 10% of the leader, MiniMax M2.1 from MiniMax is the cheapest, at $0.30 per million input tokens with a score of 0.479.

How recent are the Terminal-Bench leaderboard results?

The Terminal-Bench leaderboard was last updated in October 2026 and currently includes 25 evaluated models.