AA-Briefcase
Progress Over Time
Interactive timeline showing model performance evolution on AA-Briefcase
AA-Briefcase Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | xAI | — | 500K | $2.00 / $6.00 | ||
| 2 | Moonshot AI | 2.8T | 1.0M | $2.85 / $14.25 | ||
| 3 | Mistral AI | 1.1T | 1.0M | $0.68 / $2.09 | ||
| 4 | Thinking Machines Lab | 276B | 524K | $0.30 / $1.20 |
Sub-benchmarks
What is AA-Briefcase?
AA-Briefcase is an Artificial Analysis evaluation of AI systems on professional knowledge-work tasks, reported as an Elo score.
AA-Briefcase is a text benchmark evaluating models on productivity, reasoning, and agents tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–3000 scale. The current average is 1358.8, with the leader at 1577.0.
Compare leaders on the best AI for productivity, best AI for reasoning and best AI for agents leaderboards.
Current leaders
Grok 4.6 from xAI currently leads the AA-Briefcase leaderboard with a score of 1577.000 across 4 evaluated AI models.
FAQ
Common questions about the AA-Briefcase benchmark and leaderboard.