The AI arena is free today

Open Superagent

NL2Repo

Progress Over Time

Interactive timeline showing model performance evolution on NL2Repo

State-of-the-art frontier
Open
Proprietary

NL2Repo Leaderboard

22 models
ContextCostLicense
11.6T1.0M$0.43 / $0.87
2770B
3
Zhipu AI
Zhipu AI
753B1.0M$1.40 / $4.40
41.0M$0.22 / $0.66
5320B1.0M$0.15 / $0.50
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
7304B1.0M$0.09 / $0.18
8
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B1.0M$0.15 / $0.47
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
12
ByteDance
ByteDance
13
Tencent
Tencent
295B
14
15
Zhipu AI
Zhipu AI
754B200K$1.40 / $4.40
16
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K
17
MiniMax
MiniMax
428B1.0M$0.30 / $1.20
18
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
19205K$0.30 / $1.20
20
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
21
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.60 / $3.60
22
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
Notice missing or incorrect data?
About this benchmark

What is NL2Repo?

NL2Repo evaluates long-horizon coding capabilities including repository-level understanding, where models must generate or modify code across entire repositories from natural language specifications.

NL2Repo is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 22 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.6.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

DeepSeek-V4-Pro-0813 from DeepSeek currently leads the NL2Repo leaderboard with a score of 0.615 across 22 evaluated AI models.

1DeepSeek-V4-Pro-0813DeepSeek61.5%
2Hy4 previewTencent58.9%
3GLM-5.3Zhipu AI58.0%

FAQ

Common questions about the NL2Repo benchmark and leaderboard.

What is the NL2Repo benchmark?

NL2Repo evaluates long-horizon coding capabilities including repository-level understanding, where models must generate or modify code across entire repositories from natural language specifications.

What is the NL2Repo leaderboard?

The NL2Repo leaderboard ranks 22 AI models based on their performance on this benchmark. Currently, DeepSeek-V4-Pro-0813 by DeepSeek leads with a score of 0.615. The average score across all models is 0.474.

What is the highest NL2Repo score?

The highest NL2Repo score is 0.615, achieved by DeepSeek-V4-Pro-0813 from DeepSeek.

How many models are evaluated on NL2Repo?

22 models have been evaluated on the NL2Repo benchmark, with 0 verified results and 22 self-reported results.

What categories does NL2Repo cover?

NL2Repo is categorized under agents and code. The benchmark evaluates text models.

What is the best open-source model on NL2Repo?

DeepSeek-V4-Pro-0813 by DeepSeek is the top-ranked open-source model on NL2Repo, with a score of 0.615 (rank #1).

Which model offers the best value on NL2Repo?

Among models scoring within 10% of the leader, GLM-5.3-Flash from Zhipu AI is the cheapest, at $0.15 per million input tokens with a score of 0.563.

How recent are the NL2Repo leaderboard results?

The NL2Repo leaderboard was last updated in September 2026 and currently includes 22 evaluated models.