NL2Repo
Progress Over Time
Interactive timeline showing model performance evolution on NL2Repo
NL2Repo Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | DeepSeek | 1.6T | 1.0M | $0.43 / $0.87 | ||
| 2 | Tencent | 770B | — | — | ||
| 3 | Zhipu AI | 753B | 1.0M | $1.40 / $4.40 | ||
| 4 | DeepSeek | — | 1.0M | $0.22 / $0.66 | ||
| 5 | Zhipu AI | 320B | 1.0M | $0.15 / $0.50 | ||
| 6 | Alibaba Cloud / Qwen Team | 2.4T | 1.0M | $1.65 / $4.95 | ||
| 7 | DeepSeek | 304B | 1.0M | $0.09 / $0.18 | ||
| 8 | Zhipu AI | 753B | 1.0M | $0.95 / $3.00 | ||
| 9 | Alibaba Cloud / Qwen Team | 125B | 1.0M | $0.15 / $0.47 | ||
| 9 | Alibaba Cloud / Qwen Team | 125B | — | — | ||
| 11 | Alibaba Cloud / Qwen Team | — | 1.0M | $1.25 / $3.75 | ||
| 12 | ByteDance | — | — | — | ||
| 13 | Tencent | 295B | — | — | ||
| 14 | ByteDance | — | — | — | ||
| 15 | Zhipu AI | 754B | 200K | $1.40 / $4.40 | ||
| 16 | Alibaba Cloud / Qwen Team | 28B | 262K | — | ||
| 17 | MiniMax | 428B | 1.0M | $0.30 / $1.20 | ||
| 18 | Alibaba Cloud / Qwen Team | — | — | — | ||
| 19 | MiniMax | — | 205K | $0.30 / $1.20 | ||
| 20 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 21 | Alibaba Cloud / Qwen Team | 28B | 262K | $0.60 / $3.60 | ||
| 22 | Alibaba Cloud / Qwen Team | 35B | — | — |
What is NL2Repo?
NL2Repo evaluates long-horizon coding capabilities including repository-level understanding, where models must generate or modify code across entire repositories from natural language specifications.
NL2Repo is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 22 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.6.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
DeepSeek-V4-Pro-0813 from DeepSeek currently leads the NL2Repo leaderboard with a score of 0.615 across 22 evaluated AI models.
FAQ
Common questions about the NL2Repo benchmark and leaderboard.