Program Bench
Progress Over Time
Interactive timeline showing model performance evolution on Program Bench
Program Bench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $4.00 / $20.00 | ||
| 2 | Anthropic | — | 1.0M | $0.10 / $0.50 | ||
| 3 | Anthropic | — | 1.0M | $2.00 / $10.00 | ||
| 4 | Moonshot AI | 2.8T | 1.0M | $2.85 / $14.25 | ||
| 5 | Zhipu AI | 753B | 1.0M | $0.75 / $2.40 | ||
| 6 | Moonshot AI | 1.0T | 262K | $0.68 / $3.40 | ||
| 7 | ByteDance | — | — | — | ||
| 8 | ByteDance | — | — | — | ||
| 9 | Xiaomi | 1.0T | 1.0M | $0.43 / $0.87 | ||
| 10 | Xiaomi | 309B | 1.0M | $0.14 / $0.28 | ||
| 11 | DeepSeek | 763B | 1.0M | $0.22 / $0.66 | ||
| 12 | Zhipu AI | 753B | 1.0M | $1.20 / $4.00 | ||
| 13 | Tencent | 770B | — | — |
What is Program Bench?
Program Bench evaluates code-generation agents by asking them to recreate a program's behavior from only a compiled binary and documentation. It spans 200 tasks from small CLI tools to large systems such as FFmpeg and SQLite, with submissions judged against more than 248,000 fuzz-generated behavioral tests.
Program Bench is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 13 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.9.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
Claude Opus 5.5 from Anthropic currently leads the Program Bench leaderboard with a score of 0.912 across 13 evaluated AI models.
FAQ
Common questions about the Program Bench benchmark and leaderboard.