Legal Agent Benchmark
Progress Over Time
Interactive timeline showing model performance evolution on Legal Agent Benchmark
Legal Agent Benchmark Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 2 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 3 | Anthropic | — | 1.0M | $3.00 / $15.00 | ||
| 4 | Anthropic | — | 200K | $3.00 / $15.00 | ||
| 5 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 6 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 7 | Google | — | 1.0M | $1.50 / $9.00 | ||
| 8 | OpenAI | — | 1.0M | $2.50 / $15.00 | ||
| 9 | Google | — | 1.0M | $2.50 / $15.00 | ||
| 9 | Google | — | 1.0M | $0.50 / $3.00 | ||
| 9 | OpenAI | — | 400K | $0.75 / $4.50 | ||
| 9 | Google | — | 1.0M | $0.25 / $1.50 |
What is Legal Agent Benchmark?
The Legal Agent Benchmark (LAB) is Harvey's open-source benchmark for evaluating AI agents on complex, long-horizon legal work. Tasks are scored under an all-pass standard against expert-curated rubrics, where a task passes only if every required rubric criterion (facts, conclusions, citations, structure, and analytical moves) passes.
Legal Agent Benchmark is a text benchmark evaluating models on legal, reasoning, and agents tasks. LLM Stats tracks 12 models on this benchmark, scored on a 0–1 scale. The current average is 0.0, with the leader at 0.1.
Compare leaders on the best AI for legal, best AI for reasoning and best AI for agents leaderboards.
Current leaders
Claude Fable 5 from Anthropic currently leads the Legal Agent Benchmark leaderboard with a score of 0.133 across 12 evaluated AI models.
FAQ
Common questions about the Legal Agent Benchmark benchmark and leaderboard.