ExploitBench
Progress Over Time
Interactive timeline showing model performance evolution on ExploitBench
ExploitBench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 2 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 3 | OpenAI | — | 1.1M | $2.50 / $15.00 | ||
| 4 | OpenAI | — | 1.1M | $1.00 / $6.00 |
What is ExploitBench?
ExploitBench is a cybersecurity benchmark that evaluates a model's ability to discover and exploit software vulnerabilities, reported as the fraction of challenges where the model captures the target (Cap%).
ExploitBench is a text benchmark evaluating models on safety, agents, and code tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.8.
Compare leaders on the best AI for safety, best AI for agents and best AI for code leaderboards.
Current leaders
Claude Fable 5 from Anthropic currently leads the ExploitBench leaderboard with a score of 0.780 across 4 evaluated AI models.
FAQ
Common questions about the ExploitBench benchmark and leaderboard.