GDPval-Rubrics
Progress Over Time
Interactive timeline showing model performance evolution on GDPval-Rubrics
GDPval-Rubrics Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | MiniMax | — | 1.0M | $0.30 / $1.20 |
What is GDPval-Rubrics?
GDPval-Rubrics evaluates AI model performance on economically valuable knowledge work tasks drawn from the public GDPval dataset. It uses pointwise scoring based on public rubrics, with the environment aligned to the GDPval-AA scaffolding.
GDPval-Rubrics is a text benchmark evaluating models on legal, reasoning, finance, general, and agents tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.7.
Compare leaders on the best AI for legal, best AI for reasoning, best AI for finance, best AI for general and best AI for agents leaderboards.
Current leaders
MiniMax M3 from MiniMax currently leads the GDPval-Rubrics leaderboard with a score of 0.748 across 1 evaluated AI models.
FAQ
Common questions about the GDPval-Rubrics benchmark and leaderboard.