AA-Briefcase v1.1
Progress Over Time
Interactive timeline showing model performance evolution on AA-Briefcase v1.1
AA-Briefcase v1.1 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $4.00 / $20.00 | ||
| 2 | Anthropic | — | 1.0M | $2.00 / $10.00 | ||
| 3 | xAI | — | 500K | $2.00 / $6.00 | ||
| 4 | Anthropic | — | 1.0M | $0.10 / $0.50 | ||
| 5 | Upstage | 35B | 524K | $0.10 / $0.40 |
What is AA-Briefcase v1.1?
AA-Briefcase v1.1 evaluates multi-hour professional office work and reports Elo scores. Kept separate from unversioned AA-Briefcase results.
AA-Briefcase v1.1 is a text benchmark evaluating models on productivity, reasoning, and agents tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–3000 scale. The current average is 1548.0, with the leader at 1822.0.
Compare leaders on the best AI for productivity, best AI for reasoning and best AI for agents leaderboards.
Current leaders
Claude Opus 5.5 from Anthropic currently leads the AA-Briefcase v1.1 leaderboard with a score of 1822.000 across 5 evaluated AI models.
FAQ
Common questions about the AA-Briefcase v1.1 benchmark and leaderboard.