Kimi Code Bench v2
Progress Over Time
Interactive timeline showing model performance evolution on Kimi Code Bench v2
Kimi Code Bench v2 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Kimi K3New Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 2 | Moonshot AI | 1.0T | 262K | $0.74 / $3.50 |
What is Kimi Code Bench v2?
Kimi Code Bench v2 is Moonshot AI's in-house benchmark for evaluating coding agents on realistic software engineering tasks across 10+ mainstream programming languages and a production tech stack spanning backend services, infrastructure, performance engineering, systems programming, security, frontend development, and ML/data engineering.
Kimi Code Bench v2 is a text benchmark evaluating models on agents and coding tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.7.
Compare leaders on the best AI for agents and best AI for coding leaderboards.
Current leaders
Kimi K3 from Moonshot AI currently leads the Kimi Code Bench v2 leaderboard with a score of 0.729 across 2 evaluated AI models.
FAQ
Common questions about the Kimi Code Bench v2 benchmark and leaderboard.