The AI arena is free today

Open Superagent

OfficeQA

Progress Over Time

Interactive timeline showing model performance evolution on OfficeQA

State-of-the-art frontier
Open
Proprietary

OfficeQA Leaderboard

3 models
ContextCostLicense
1—1.0M$4.00 / $20.00
2—1.0M$2.00 / $10.00
3
Anthropic
Anthropic
—1.0M$0.10 / $0.50
Notice missing or incorrect data?
About this benchmark

What is OfficeQA?

OfficeQA is a Databricks public benchmark that evaluates end-to-end grounded reasoning over a large corpus of historical U.S. Treasury Bulletin documents. Models must locate relevant tables across the corpus and perform precise numerical reasoning over them.

OfficeQA is a text benchmark evaluating models on reasoning, general, and agents tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.

Compare leaders on the best AI for reasoning, best AI for general and best AI for agents leaderboards.

Current leaders

Claude Opus 5.5 from Anthropic currently leads the OfficeQA leaderboard with a score of 0.789 across 3 evaluated AI models.

1Claude Opus 5.5Anthropic78.9%
2Claude Sonnet 5.5Anthropic76.9%
3Claude Haiku 5.5Anthropic73.5%

FAQ

Common questions about the OfficeQA benchmark and leaderboard.

What is the OfficeQA benchmark?

OfficeQA is a Databricks public benchmark that evaluates end-to-end grounded reasoning over a large corpus of historical U.S. Treasury Bulletin documents. Models must locate relevant tables across the corpus and perform precise numerical reasoning over them.

What is the OfficeQA leaderboard?

The OfficeQA leaderboard ranks 3 AI models based on their performance on this benchmark. Currently, Claude Opus 5.5 by Anthropic leads with a score of 0.789. The average score across all models is 0.764.

What is the highest OfficeQA score?

The highest OfficeQA score is 0.789, achieved by Claude Opus 5.5 from Anthropic.

How many models are evaluated on OfficeQA?

3 models have been evaluated on the OfficeQA benchmark, with 0 verified results and 3 self-reported results.

What categories does OfficeQA cover?

OfficeQA is categorized under reasoning, general, and agents. The benchmark evaluates text models.

Which model offers the best value on OfficeQA?

Among models scoring within 10% of the leader, Claude Haiku 5.5 from Anthropic is the cheapest, at $0.10 per million input tokens with a score of 0.735.

How recent are the OfficeQA leaderboard results?

The OfficeQA leaderboard was last updated in October 2026 and currently includes 3 evaluated models.