SWE Atlas - Codebase QnA
Progress Over Time
Interactive timeline showing model performance evolution on SWE Atlas - Codebase QnA
SWE Atlas - Codebase QnA Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Mistral AI | 1.1T | 1.0M | $0.68 / $2.09 | ||
| 1 | Meta | — | 1.0M | $0.10 / $0.20 | ||
| 3 | InclusionAI | 560B | — | — | ||
| 4 | Poolside | 118B | 1.0M | $0.10 / $0.20 | ||
| 5 | MiniMax | 428B | 1.0M | $0.28 / $1.10 |
What is SWE Atlas - Codebase QnA?
SWE Atlas - Codebase QnA evaluates a model's ability to answer questions about real codebases, measuring repository-level comprehension and the ability to reason about code structure, behavior, and intent across an entire project.
SWE Atlas - Codebase QnA is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.6.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
Mistral Large 4 from Mistral AI currently leads the SWE Atlas - Codebase QnA leaderboard with a score of 0.594 across 5 evaluated AI models.
FAQ
Common questions about the SWE Atlas - Codebase QnA benchmark and leaderboard.