SWE-Marathon
Progress Over Time
Interactive timeline showing model performance evolution on SWE-Marathon
SWE-Marathon Leaderboard
What is SWE-Marathon?
SWE-Marathon is an ultra-long-horizon software engineering benchmark covering tasks such as building compilers, optimizing kernels, and developing production-grade services. It measures whether agents can sustain quality across extremely long engineering trajectories.
SWE-Marathon is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.4.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
Kimi K3 from Moonshot AI currently leads the SWE-Marathon leaderboard with a score of 0.420 across 3 evaluated AI models.
FAQ
Common questions about the SWE-Marathon benchmark and leaderboard.