APEX-SWE
Progress Over Time
Interactive timeline showing model performance evolution on APEX-SWE
State-of-the-art frontier
Open
Proprietary
APEX-SWE Leaderboard
1 models
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | xAI | — | 500K | $2.00 / $6.00 |
Notice missing or incorrect data?
What is APEX-SWE?
APEX-SWE evaluates AI agents on software engineering tasks requiring multi-step coding, debugging, and verification.
APEX-SWE is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.6.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
Grok 4.6 from xAI currently leads the APEX-SWE leaderboard with a score of 0.564 across 1 evaluated AI models.
FAQ
Common questions about the APEX-SWE benchmark and leaderboard.