The AI arena is free today

Open Superagent

AA-Briefcase v1.1

Implementation

Progress Over Time

Interactive timeline showing model performance evolution on AA-Briefcase v1.1

State-of-the-art frontier
Open
Proprietary

AA-Briefcase v1.1 Leaderboard

5 models
ContextCostLicense
1—1.0M$4.00 / $20.00
2—1.0M$2.00 / $10.00
3—500K$2.00 / $6.00
4
Anthropic
Anthropic
—1.0M$0.10 / $0.50
535B524K$0.10 / $0.40
Notice missing or incorrect data?
About this benchmark

What is AA-Briefcase v1.1?

AA-Briefcase v1.1 evaluates multi-hour professional office work and reports Elo scores. Kept separate from unversioned AA-Briefcase results.

AA-Briefcase v1.1 is a text benchmark evaluating models on productivity, reasoning, and agents tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–3000 scale. The current average is 1548.0, with the leader at 1822.0.

Compare leaders on the best AI for productivity, best AI for reasoning and best AI for agents leaderboards.

Current leaders

Claude Opus 5.5 from Anthropic currently leads the AA-Briefcase v1.1 leaderboard with a score of 1822.000 across 5 evaluated AI models.

1Claude Opus 5.5Anthropic1822.000
2Claude Sonnet 5.5Anthropic1811.000
3Grok 4.7xAI1657.000

FAQ

Common questions about the AA-Briefcase v1.1 benchmark and leaderboard.

What is the AA-Briefcase v1.1 benchmark?

AA-Briefcase v1.1 evaluates multi-hour professional office work and reports Elo scores. Kept separate from unversioned AA-Briefcase results.

What is the AA-Briefcase v1.1 leaderboard?

The AA-Briefcase v1.1 leaderboard ranks 5 AI models based on their performance on this benchmark. Currently, Claude Opus 5.5 by Anthropic leads with a score of 1822.000. The average score across all models is 1548.000.

What is the highest AA-Briefcase v1.1 score?

The highest AA-Briefcase v1.1 score is 1822.000, achieved by Claude Opus 5.5 from Anthropic.

How many models are evaluated on AA-Briefcase v1.1?

5 models have been evaluated on the AA-Briefcase v1.1 benchmark, with 0 verified results and 4 self-reported results.

Where can I find the AA-Briefcase v1.1 dataset?

The AA-Briefcase v1.1 dataset is available at https://x.ai/news/grok-4-7.

What categories does AA-Briefcase v1.1 cover?

AA-Briefcase v1.1 is categorized under productivity, reasoning, and agents. The benchmark evaluates text models.

What's the difference between AA-Briefcase v1.1 and AA-Briefcase?

AA-Briefcase v1.1 is a variant of AA-Briefcase. See the AA-Briefcase leaderboard for the broader benchmark and per-model comparison.

Which model offers the best value on AA-Briefcase v1.1?

Among models scoring within 10% of the leader, Claude Sonnet 5.5 from Anthropic is the cheapest, at $2.00 per million input tokens with a score of 1811.000.

How recent are the AA-Briefcase v1.1 leaderboard results?

The AA-Briefcase v1.1 leaderboard was last updated in October 2026 and currently includes 5 evaluated models.