The AI arena is free today

Open Superagent

Terminal-Bench 4.0

Implementation

Progress Over Time

Interactive timeline showing model performance evolution on Terminal-Bench 4.0

State-of-the-art frontier
Open
Proprietary

Terminal-Bench 4.0 Leaderboard

13 models
ContextCostLicense
1
OpenAI
OpenAI
1.1M$10.00 / $50.00
2
Anthropic
Anthropic
1.0M$10.00 / $50.00
3
Anthropic
Anthropic
1.0M$5.00 / $25.00
41.0M$10.00 / $50.00
5
Zhipu AI
Zhipu AI
753B1.0M$1.40 / $4.40
61.1M$5.00 / $30.00
71.0M$5.00 / $25.00
81.1M$2.00 / $12.00
9500K$2.00 / $6.00
101.0M$0.75 / $3.75
111.1M$0.20 / $1.20
121.0M$2.00 / $10.00
12500K$2.00 / $6.00
Notice missing or incorrect data?
About this benchmark

What is Terminal-Bench 4.0?

Terminal-Bench 4.0 evaluates AI agents on real-world work in terminal and command-line containerized environments across 66 tasks, with emphasis on science-adjacent and frontier engineering problems. Relative to earlier Terminal-Bench releases, 4.0 increases timeouts and adaptively raises RAM/CPU on selected tasks to reduce harness and resource confounds.

Terminal-Bench 4.0 is a text benchmark evaluating models on reasoning, agents, code, and tool calling tasks. LLM Stats tracks 13 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.6.

Compare leaders on the best AI for reasoning, best AI for agents, best AI for code and best AI for tool calling leaderboards.

Current leaders

GPT-6 Astra from OpenAI currently leads the Terminal-Bench 4.0 leaderboard with a score of 0.577 across 13 evaluated AI models.

1GPT-6 AstraOpenAI57.7%
2Claude Fable 5.1Anthropic55.8%
3Claude Opus 5Anthropic51.8%

FAQ

Common questions about the Terminal-Bench 4.0 benchmark and leaderboard.

What is the Terminal-Bench 4.0 benchmark?

Terminal-Bench 4.0 evaluates AI agents on real-world work in terminal and command-line containerized environments across 66 tasks, with emphasis on science-adjacent and frontier engineering problems. Relative to earlier Terminal-Bench releases, 4.0 increases timeouts and adaptively raises RAM/CPU on selected tasks to reduce harness and resource confounds.

What is the Terminal-Bench 4.0 leaderboard?

The Terminal-Bench 4.0 leaderboard ranks 13 AI models based on their performance on this benchmark. Currently, GPT-6 Astra by OpenAI leads with a score of 0.577. The average score across all models is 0.320.

What is the highest Terminal-Bench 4.0 score?

The highest Terminal-Bench 4.0 score is 0.577, achieved by GPT-6 Astra from OpenAI.

How many models are evaluated on Terminal-Bench 4.0?

13 models have been evaluated on the Terminal-Bench 4.0 benchmark, with 0 verified results and 3 self-reported results.

Where can I find the Terminal-Bench 4.0 dataset?

The Terminal-Bench 4.0 dataset is available at https://github.com/laude-institute/terminal-bench.

What categories does Terminal-Bench 4.0 cover?

Terminal-Bench 4.0 is categorized under reasoning, agents, code, and tool calling. The benchmark evaluates text models.

What's the difference between Terminal-Bench 4.0 and Terminal-Bench?

Terminal-Bench 4.0 is a variant of Terminal-Bench. See the Terminal-Bench leaderboard for the broader benchmark and per-model comparison.

Which model offers the best value on Terminal-Bench 4.0?

Among models scoring within 10% of the leader, GPT-6 Astra from OpenAI is the cheapest, at $10.00 per million input tokens with a score of 0.577.

How recent are the Terminal-Bench 4.0 leaderboard results?

The Terminal-Bench 4.0 leaderboard was last updated in September 2026 and currently includes 13 evaluated models.