Legal Agent Benchmark
Progress Over Time
Interactive timeline showing model performance evolution on Legal Agent Benchmark
Legal Agent Benchmark Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 2 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 3 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 4 | Anthropic | — | 1.0M | $3.00 / $15.00 | ||
| 5 | Anthropic | — | 200K | $3.00 / $15.00 | ||
| 6 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 7 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 8 | Google | — | 1.0M | $1.50 / $9.00 | ||
| 9 | OpenAI | — | 1.0M | $2.50 / $15.00 | ||
| 10 | Google | — | 1.0M | $2.50 / $15.00 | ||
| 10 | Google | — | 1.0M | $0.50 / $3.00 | ||
| 10 | OpenAI | — | 400K | $0.75 / $4.50 | ||
| 10 | Google | — | 1.0M | $0.25 / $1.50 |
What is Legal Agent Benchmark?
The Legal Agent Benchmark (LAB) is Harvey's open-source benchmark for evaluating AI agents on complex, long-horizon legal work. Tasks are scored under an all-pass standard against expert-curated rubrics, where a task passes only if every required rubric criterion (facts, conclusions, citations, structure, and analytical moves) passes.
Legal Agent Benchmark is a text benchmark evaluating models on reasoning, legal, and agents tasks. LLM Stats tracks 13 models on this benchmark, scored on a 0–1 scale. The current average is 0.0, with the leader at 0.1.
Compare leaders on the best AI for reasoning, best AI for legal and best AI for agents leaderboards.
Current leaders
Claude Fable 5 from Anthropic currently leads the Legal Agent Benchmark leaderboard with a score of 0.133 across 13 evaluated AI models.
FAQ
Common questions about the Legal Agent Benchmark benchmark and leaderboard.