Terminal-Bench-Science 0.1
Progress Over Time
Interactive timeline showing model performance evolution on Terminal-Bench-Science 0.1
Terminal-Bench-Science 0.1 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | GPT-6 AstraNew OpenAI | — | 1.1M | $10.00 / $50.00 | ||
| 2 | Anthropic | — | 1.0M | $10.00 / $50.00 |
What is Terminal-Bench-Science 0.1?
Terminal-Bench-Science 0.1 is a Stanford-led community benchmark of 70 tasks drawn from scientific research workflows across the life, physical, earth, mathematical, and engineering sciences. Tasks are authored and reviewed by scientists and researchers; scores are typically reported as accuracy with relatively large standard error due to strongly bimodal per-task outcomes.
Terminal-Bench-Science 0.1 is a text benchmark evaluating models on reasoning, general, agents, code, and tool calling tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.6.
Compare leaders on the best AI for reasoning, best AI for general, best AI for agents, best AI for code and best AI for tool calling leaderboards.
Current leaders
GPT-6 Astra from OpenAI currently leads the Terminal-Bench-Science 0.1 leaderboard with a score of 0.646 across 2 evaluated AI models.
FAQ
Common questions about the Terminal-Bench-Science 0.1 benchmark and leaderboard.