The AI arena is free today

Open Superagent

Harvey LAB-AA

Progress Over Time

Interactive timeline showing model performance evolution on Harvey LAB-AA

State-of-the-art frontier
Open
Proprietary

Harvey LAB-AA Leaderboard

1 models
ContextCostLicense
11.0M$0.75 / $3.75
Notice missing or incorrect data?
About this benchmark

What is Harvey LAB-AA?

Harvey LAB-AA evaluates model performance on complex legal workflows.

Harvey LAB-AA is a text benchmark evaluating models on knowledge and agents tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 0.9.

Compare leaders on the best AI for knowledge and best AI for agents leaderboards.

Current leaders

Gemini 3.7 Flash from Google currently leads the Harvey LAB-AA leaderboard with a score of 0.907 across 1 evaluated AI models.

1Gemini 3.7 FlashGoogle90.7%

FAQ

Common questions about the Harvey LAB-AA benchmark and leaderboard.

What is the Harvey LAB-AA benchmark?

Harvey LAB-AA evaluates model performance on complex legal workflows.

What is the Harvey LAB-AA leaderboard?

The Harvey LAB-AA leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Gemini 3.7 Flash by Google leads with a score of 0.907. The average score across all models is 0.907.

What is the highest Harvey LAB-AA score?

The highest Harvey LAB-AA score is 0.907, achieved by Gemini 3.7 Flash from Google.

How many models are evaluated on Harvey LAB-AA?

1 models have been evaluated on the Harvey LAB-AA benchmark, with 0 verified results and 1 self-reported results.

What categories does Harvey LAB-AA cover?

Harvey LAB-AA is categorized under knowledge and agents. The benchmark evaluates text models.

Which model offers the best value on Harvey LAB-AA?

Among models scoring within 10% of the leader, Gemini 3.7 Flash from Google is the cheapest, at $0.75 per million input tokens with a score of 0.907.

How recent are the Harvey LAB-AA leaderboard results?

The Harvey LAB-AA leaderboard was last updated in August 2026 and currently includes 1 evaluated models.