The AI arena is free today

Open Superagent

SWE Atlas - Codebase QnA

Progress Over Time

Interactive timeline showing model performance evolution on SWE Atlas - Codebase QnA

State-of-the-art frontier
Open
Proprietary

SWE Atlas - Codebase QnA Leaderboard

5 models
ContextCostLicense
1
Mistral AI
Mistral AI
1.1T1.0M$0.68 / $2.09
1—1.0M$0.10 / $0.20
3
InclusionAI
InclusionAI
560B——
4
Poolside
Poolside
118B1.0M$0.10 / $0.20
5
MiniMax
MiniMax
428B1.0M$0.28 / $1.10
Notice missing or incorrect data?
About this benchmark

What is SWE Atlas - Codebase QnA?

SWE Atlas - Codebase QnA evaluates a model's ability to answer questions about real codebases, measuring repository-level comprehension and the ability to reason about code structure, behavior, and intent across an entire project.

SWE Atlas - Codebase QnA is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.6.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

Mistral Large 4 from Mistral AI currently leads the SWE Atlas - Codebase QnA leaderboard with a score of 0.594 across 5 evaluated AI models.

1Mistral Large 4Mistral AI59.4%
1Muse Spark 1.3Meta59.4%
3Ling 3.1 FlashInclusionAI55.9%
OSSLaguna S 2.1#4 open-weight46.2%

FAQ

Common questions about the SWE Atlas - Codebase QnA benchmark and leaderboard.

What is the SWE Atlas - Codebase QnA benchmark?

SWE Atlas - Codebase QnA evaluates a model's ability to answer questions about real codebases, measuring repository-level comprehension and the ability to reason about code structure, behavior, and intent across an entire project.

What is the SWE Atlas - Codebase QnA leaderboard?

The SWE Atlas - Codebase QnA leaderboard ranks 5 AI models based on their performance on this benchmark. Currently, Mistral Large 4 by Mistral AI leads with a score of 0.594. The average score across all models is 0.518.

What is the highest SWE Atlas - Codebase QnA score?

The highest SWE Atlas - Codebase QnA score is 0.594, achieved by Mistral Large 4 from Mistral AI.

How many models are evaluated on SWE Atlas - Codebase QnA?

5 models have been evaluated on the SWE Atlas - Codebase QnA benchmark, with 0 verified results and 5 self-reported results.

What categories does SWE Atlas - Codebase QnA cover?

SWE Atlas - Codebase QnA is categorized under agents and code. The benchmark evaluates text models.

What is the best open-source model on SWE Atlas - Codebase QnA?

Laguna S 2.1 by Poolside is the top-ranked open-source model on SWE Atlas - Codebase QnA, with a score of 0.462 (rank #4).

Which model offers the best value on SWE Atlas - Codebase QnA?

Among models scoring within 10% of the leader, Muse Spark 1.3 from Meta is the cheapest, at $0.10 per million input tokens with a score of 0.594.

How recent are the SWE Atlas - Codebase QnA leaderboard results?

The SWE Atlas - Codebase QnA leaderboard was last updated in October 2026 and currently includes 5 evaluated models.