SWE-Marathon

Progress Over Time

Interactive timeline showing model performance evolution on SWE-Marathon

State-of-the-art frontier
Open
Proprietary

SWE-Marathon Leaderboard

3 models
ContextCostLicense
1
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
2500K$2.00 / $6.00
3
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
Notice missing or incorrect data?
About this benchmark

What is SWE-Marathon?

SWE-Marathon is an ultra-long-horizon software engineering benchmark covering tasks such as building compilers, optimizing kernels, and developing production-grade services. It measures whether agents can sustain quality across extremely long engineering trajectories.

SWE-Marathon is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.4.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

Kimi K3 from Moonshot AI currently leads the SWE-Marathon leaderboard with a score of 0.420 across 3 evaluated AI models.

1Kimi K3Moonshot AI42.0%
2Grok 4.5xAI29.0%
3GLM-5.2Zhipu AI13.0%

FAQ

Common questions about the SWE-Marathon benchmark and leaderboard.

What is the SWE-Marathon benchmark?

SWE-Marathon is an ultra-long-horizon software engineering benchmark covering tasks such as building compilers, optimizing kernels, and developing production-grade services. It measures whether agents can sustain quality across extremely long engineering trajectories.

What is the SWE-Marathon leaderboard?

The SWE-Marathon leaderboard ranks 3 AI models based on their performance on this benchmark. Currently, Kimi K3 by Moonshot AI leads with a score of 0.420. The average score across all models is 0.280.

What is the highest SWE-Marathon score?

The highest SWE-Marathon score is 0.420, achieved by Kimi K3 from Moonshot AI.

How many models are evaluated on SWE-Marathon?

3 models have been evaluated on the SWE-Marathon benchmark, with 0 verified results and 3 self-reported results.

What categories does SWE-Marathon cover?

SWE-Marathon is categorized under agents and code. The benchmark evaluates text models.

What is the best open-source model on SWE-Marathon?

Kimi K3 by Moonshot AI is the top-ranked open-source model on SWE-Marathon, with a score of 0.420 (rank #1).

Which model offers the best value on SWE-Marathon?

Among models scoring within 10% of the leader, Kimi K3 from Moonshot AI is the cheapest, at $3.00 per million input tokens with a score of 0.420.

How recent are the SWE-Marathon leaderboard results?

The SWE-Marathon leaderboard was last updated in July 2026 and currently includes 3 evaluated models.