Terminal-Bench 2.1

Progress Over Time

Interactive timeline showing model performance evolution on Terminal-Bench 2.1

State-of-the-art frontier
Open
Proprietary

Terminal-Bench 2.1 Leaderboard

12 models
ContextCostLicense
11.1M$5.00 / $30.00
2
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
31.1M$2.50 / $15.00
41.1M$1.00 / $6.00
51.0M$10.00 / $50.00
6500K$2.00 / $6.00
7
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
8
Tencent
Tencent
295B
9
ByteDance
ByteDance
10
11
MiniMax
MiniMax
1.0M$0.30 / $1.20
12550B
Notice missing or incorrect data?
About this benchmark

What is Terminal-Bench 2.1?

Terminal-Bench 2.1 is an updated release of the Terminal-Bench benchmark that tests AI agents' ability to operate a computer via the terminal. It evaluates how well models handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, data science workflows, and security tasks.

Terminal-Bench 2.1 is a text benchmark evaluating models on reasoning, agents, code, and tool calling tasks. LLM Stats tracks 12 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.9.

Compare leaders on the best AI for reasoning, best AI for agents, best AI for code and best AI for tool calling leaderboards.

Current leaders

GPT-5.6 Sol from OpenAI currently leads the Terminal-Bench 2.1 leaderboard with a score of 0.888 across 12 evaluated AI models.

1GPT-5.6 SolOpenAI88.8%
2Kimi K3Moonshot AI88.3%
3GPT-5.6 TerraOpenAI87.4%

FAQ

Common questions about the Terminal-Bench 2.1 benchmark and leaderboard.

What is the Terminal-Bench 2.1 benchmark?

Terminal-Bench 2.1 is an updated release of the Terminal-Bench benchmark that tests AI agents' ability to operate a computer via the terminal. It evaluates how well models handle real-world, end-to-end tasks autonomously, including compiling code, training models, setting up servers, system administration, data science workflows, and security tasks.

What is the Terminal-Bench 2.1 leaderboard?

The Terminal-Bench 2.1 leaderboard ranks 12 AI models based on their performance on this benchmark. Currently, GPT-5.6 Sol by OpenAI leads with a score of 0.888. The average score across all models is 0.777.

What is the highest Terminal-Bench 2.1 score?

The highest Terminal-Bench 2.1 score is 0.888, achieved by GPT-5.6 Sol from OpenAI.

How many models are evaluated on Terminal-Bench 2.1?

12 models have been evaluated on the Terminal-Bench 2.1 benchmark, with 0 verified results and 12 self-reported results.

What categories does Terminal-Bench 2.1 cover?

Terminal-Bench 2.1 is categorized under reasoning, agents, code, and tool calling. The benchmark evaluates text models.

What's the difference between Terminal-Bench 2.1 and Terminal-Bench?

Terminal-Bench 2.1 is a variant of Terminal-Bench. See the Terminal-Bench leaderboard for the broader benchmark and per-model comparison.

What is the best open-source model on Terminal-Bench 2.1?

Kimi K3 by Moonshot AI is the top-ranked open-source model on Terminal-Bench 2.1, with a score of 0.883 (rank #2).

Which model offers the best value on Terminal-Bench 2.1?

Among models scoring within 10% of the leader, GLM-5.2 from Zhipu AI is the cheapest, at $0.95 per million input tokens with a score of 0.827.

How recent are the Terminal-Bench 2.1 leaderboard results?

The Terminal-Bench 2.1 leaderboard was last updated in July 2026 and currently includes 12 evaluated models.