The AI arena is free today

Open Superagent

SWT-Bench

Progress Over Time

Interactive timeline showing model performance evolution on SWT-Bench

State-of-the-art frontier
Open
Proprietary

SWT-Bench Leaderboard

1 models
ContextCostLicense
1230B1.0M$0.30 / $1.20
Notice missing or incorrect data?
About this benchmark

What is SWT-Bench?

Software Test Benchmark evaluating LLM ability to write tests for software repositories

SWT-Bench is a text benchmark evaluating models on code tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.7.

Compare leaders on the best AI for code leaderboards.

Current leaders

MiniMax M2.1 from MiniMax currently leads the SWT-Bench leaderboard with a score of 0.693 across 1 evaluated AI models.

1MiniMax M2.1MiniMax69.3%

FAQ

Common questions about the SWT-Bench benchmark and leaderboard.

What is the SWT-Bench benchmark?

Software Test Benchmark evaluating LLM ability to write tests for software repositories

What is the SWT-Bench leaderboard?

The SWT-Bench leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, MiniMax M2.1 by MiniMax leads with a score of 0.693. The average score across all models is 0.693.

What is the highest SWT-Bench score?

The highest SWT-Bench score is 0.693, achieved by MiniMax M2.1 from MiniMax.

How many models are evaluated on SWT-Bench?

1 models have been evaluated on the SWT-Bench benchmark, with 0 verified results and 1 self-reported results.

What categories does SWT-Bench cover?

SWT-Bench is categorized under code. The benchmark evaluates text models.

What is the best open-source model on SWT-Bench?

MiniMax M2.1 by MiniMax is the top-ranked open-source model on SWT-Bench, with a score of 0.693 (rank #1).

Which model offers the best value on SWT-Bench?

Among models scoring within 10% of the leader, MiniMax M2.1 from MiniMax is the cheapest, at $0.30 per million input tokens with a score of 0.693.

How recent are the SWT-Bench leaderboard results?

The SWT-Bench leaderboard was last updated in August 2026 and currently includes 1 evaluated models.
SWT-Bench Leaderboard