Program Bench
Progress Over Time
Interactive timeline showing model performance evolution on Program Bench
Program Bench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Kimi K3New Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 2 | Zhipu AI | 753B | 1.0M | $0.95 / $3.00 | ||
| 3 | Moonshot AI | 1.0T | 262K | $0.74 / $3.50 | ||
| 4 | ByteDance | — | — | — | ||
| 5 | ByteDance | — | — | — |
What is Program Bench?
Program Bench evaluates code-generation agents by asking them to recreate a program's behavior from only a compiled binary and documentation. It spans 200 tasks from small CLI tools to large systems such as FFmpeg and SQLite, with submissions judged against more than 248,000 fuzz-generated behavioral tests.
Program Bench is a text benchmark evaluating models on agents and coding tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.8.
Compare leaders on the best AI for agents and best AI for coding leaderboards.
Current leaders
Kimi K3 from Moonshot AI currently leads the Program Bench leaderboard with a score of 0.778 across 5 evaluated AI models.
FAQ
Common questions about the Program Bench benchmark and leaderboard.