GDPval-AA 2.1
Progress Over Time
Interactive timeline showing model performance evolution on GDPval-AA 2.1
GDPval-AA 2.1 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $4.00 / $20.00 | ||
| 2 | Anthropic | — | 1.0M | $2.00 / $10.00 | ||
| 3 | InclusionAI | 560B | — | — | ||
| 3 | Xiaomi | 1.0T | 1.0M | $0.43 / $0.87 | ||
| 5 | Anthropic | — | 1.0M | $0.10 / $0.50 | ||
| 6 | Upstage | 35B | 524K | $0.10 / $0.40 |
What is GDPval-AA 2.1?
Version 2.1 of the Artificial Analysis GDPval evaluation of professional knowledge work, reported as an Elo rating.
GDPval-AA 2.1 is a text benchmark evaluating models on reasoning, general, and agents tasks. LLM Stats tracks 6 models on this benchmark, scored on a 0–3000 scale. The current average is 1621.3, with the leader at 1846.0.
Compare leaders on the best AI for reasoning, best AI for general and best AI for agents leaderboards.
Current leaders
Claude Opus 5.5 from Anthropic currently leads the GDPval-AA 2.1 leaderboard with a score of 1846.000 across 6 evaluated AI models.
FAQ
Common questions about the GDPval-AA 2.1 benchmark and leaderboard.