The AI arena is free today

Open Superagent

AA-Briefcase

Implementation

Progress Over Time

Interactive timeline showing model performance evolution on AA-Briefcase

State-of-the-art frontier
Open
Proprietary

AA-Briefcase Leaderboard

4 models
ContextCostLicense
1—500K$2.00 / $6.00
2
Moonshot AI
Moonshot AI
2.8T1.0M$2.85 / $14.25
3
Mistral AI
Mistral AI
1.1T1.0M$0.68 / $2.09
4
Thinking Machines Lab
Thinking Machines Lab
276B524K$0.30 / $1.20
Notice missing or incorrect data?

Sub-benchmarks

About this benchmark

What is AA-Briefcase?

AA-Briefcase is an Artificial Analysis evaluation of AI systems on professional knowledge-work tasks, reported as an Elo score.

AA-Briefcase is a text benchmark evaluating models on productivity, reasoning, and agents tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–3000 scale. The current average is 1358.8, with the leader at 1577.0.

Compare leaders on the best AI for productivity, best AI for reasoning and best AI for agents leaderboards.

Current leaders

Grok 4.6 from xAI currently leads the AA-Briefcase leaderboard with a score of 1577.000 across 4 evaluated AI models.

1Grok 4.6xAI1577.000
2Kimi K3Moonshot AI1548.000
3Mistral Large 4Mistral AI1393.000
OSSInkling-Small#4 open-weight917.000

FAQ

Common questions about the AA-Briefcase benchmark and leaderboard.

What is the AA-Briefcase benchmark?

AA-Briefcase is an Artificial Analysis evaluation of AI systems on professional knowledge-work tasks, reported as an Elo score.

What is the AA-Briefcase leaderboard?

The AA-Briefcase leaderboard ranks 4 AI models based on their performance on this benchmark. Currently, Grok 4.6 by xAI leads with a score of 1577.000. The average score across all models is 1358.750.

What is the highest AA-Briefcase score?

The highest AA-Briefcase score is 1577.000, achieved by Grok 4.6 from xAI.

How many models are evaluated on AA-Briefcase?

4 models have been evaluated on the AA-Briefcase benchmark, with 0 verified results and 3 self-reported results.

Where can I find the AA-Briefcase dataset?

The AA-Briefcase dataset is available at https://artificialanalysis.ai/.

What categories does AA-Briefcase cover?

AA-Briefcase is categorized under productivity, reasoning, and agents. The benchmark evaluates text models.

Are there variants of AA-Briefcase?

Yes. AA-Briefcase has 1 related variant: AA-Briefcase v1.1.

What is the best open-source model on AA-Briefcase?

Inkling-Small by Thinking Machines Lab is the top-ranked open-source model on AA-Briefcase, with a score of 917.000 (rank #4).

Which model offers the best value on AA-Briefcase?

Among models scoring within 10% of the leader, Grok 4.6 from xAI is the cheapest, at $2.00 per million input tokens with a score of 1577.000.

How recent are the AA-Briefcase leaderboard results?

The AA-Briefcase leaderboard was last updated in October 2026 and currently includes 4 evaluated models.