The AI arena is free today

Open Superagent

Terminal-Bench 2.0

Progress Over Time

Interactive timeline showing model performance evolution on Terminal-Bench 2.0

State-of-the-art frontier
Open
Proprietary

Terminal-Bench 2.0 Leaderboard

53 models
ContextCostLicense
1
OpenAI
OpenAI
—1.1M$5.00 / $30.00
2———
3—1.0M$2.00 / $10.00
4—400K$1.75 / $14.00
5—1.0M$1.50 / $9.00
6
OpenAI
OpenAI
—1.0M$2.50 / $15.00
7—1.0M$5.00 / $25.00
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
———
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
—1.0M$1.25 / $3.75
10—1.0M$5.00 / $25.00
11
Zhipu AI
Zhipu AI
754B203K$1.05 / $3.50
12—1.0M$2.00 / $12.00
131.0T1.0M$0.43 / $0.87
141.6T1.0M$1.30 / $2.60
15
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
16
Xiaomi
Xiaomi
311B1.0M$0.17 / $0.34
17—1.0M$5.00 / $25.00
18—400K$1.75 / $14.00
19
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
—1.0M$0.50 / $3.00
20—400K$0.75 / $4.50
21
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.32 / $3.20
21———
23—1.0M$3.00 / $15.00
24———
251.0T——
26—205K$0.30 / $1.20
27284B1.0M$0.09 / $0.18
28284B1.0M$0.09 / $0.18
29
Zhipu AI
Zhipu AI
744B200K$1.00 / $3.20
30———
31———
32———
33
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B262K$0.45 / $3.00
34
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.10 / $0.95
35196B66K$0.10 / $0.40
36
Moonshot AI
Moonshot AI
1.0T——
37
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B262K$0.29 / $2.40
38—1.0M$0.50 / $3.00
39685B——
39685B——
39685B164K$0.26 / $0.38
42—400K$0.20 / $1.25
431.0T——
44
ByteDance
ByteDance
—256K$0.25 / $2.00
45
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.26 / $2.60
46
Zhipu AI
Zhipu AI
358B203K$0.40 / $1.75
47
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.14 / $1.00
48309B——
49
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
480B262K$0.30 / $1.00
4933B262K$0.10 / $0.20
1–50 of 53
1/2
Notice missing or incorrect data?
About this benchmark

What is Terminal-Bench 2.0?

Terminal-Bench 2.0 is an updated benchmark for testing AI agents' tool use ability to operate a computer via terminal. It evaluates how well models can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities.

Terminal-Bench 2.0 is a text benchmark evaluating models on reasoning, agents, code, and tool calling tasks. LLM Stats tracks 53 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.8.

Compare leaders on the best AI for reasoning, best AI for agents, best AI for code and best AI for tool calling leaderboards.

Current leaders

GPT-5.5 from OpenAI currently leads the Terminal-Bench 2.0 leaderboard with a score of 0.827 across 53 evaluated AI models.

1GPT-5.5OpenAI82.7%
2Claude Mythos PreviewAnthropic82.0%
3Claude Sonnet 5Anthropic80.4%
OSSGLM-5.1#11 open-weight69.0%

FAQ

Common questions about the Terminal-Bench 2.0 benchmark and leaderboard.

What is the Terminal-Bench 2.0 benchmark?

Terminal-Bench 2.0 is an updated benchmark for testing AI agents' tool use ability to operate a computer via terminal. It evaluates how well models can handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, security tasks, data science workflows, and cybersecurity vulnerabilities.

What is the Terminal-Bench 2.0 leaderboard?

The Terminal-Bench 2.0 leaderboard ranks 53 AI models based on their performance on this benchmark. Currently, GPT-5.5 by OpenAI leads with a score of 0.827. The average score across all models is 0.567.

What is the highest Terminal-Bench 2.0 score?

The highest Terminal-Bench 2.0 score is 0.827, achieved by GPT-5.5 from OpenAI.

How many models are evaluated on Terminal-Bench 2.0?

53 models have been evaluated on the Terminal-Bench 2.0 benchmark, with 0 verified results and 53 self-reported results.

What categories does Terminal-Bench 2.0 cover?

Terminal-Bench 2.0 is categorized under reasoning, agents, code, and tool calling. The benchmark evaluates text models.

What's the difference between Terminal-Bench 2.0 and Terminal-Bench?

Terminal-Bench 2.0 is a variant of Terminal-Bench. See the Terminal-Bench leaderboard for the broader benchmark and per-model comparison.

What is the best open-source model on Terminal-Bench 2.0?

GLM-5.1 by Zhipu AI is the top-ranked open-source model on Terminal-Bench 2.0, with a score of 0.690 (rank #11).

Which model offers the best value on Terminal-Bench 2.0?

Among models scoring within 10% of the leader, Gemini 3.5 Flash from Google is the cheapest, at $1.50 per million input tokens with a score of 0.762.

How is Terminal-Bench 2.0 scored?

Terminal-Bench 2.0 is scored using accuracy, reported on a 0–1 scale. Lower is better only when explicitly noted; on this leaderboard, higher scores indicate better performance.

How recent are the Terminal-Bench 2.0 leaderboard results?

The Terminal-Bench 2.0 leaderboard was last updated in October 2026 and currently includes 53 evaluated models.