OfficeQA Pro
Progress Over Time
Interactive timeline showing model performance evolution on OfficeQA Pro
OfficeQA Pro Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | ByteDance | — | — | — | ||
| 2 | ByteDance | — | — | — | ||
| 3 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 4 | Kimi K3New Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 5 | Anthropic | — | 1.0M | $3.00 / $15.00 | ||
| 6 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 7 | MiniMax | — | 1.0M | $0.30 / $1.20 |
What is OfficeQA Pro?
OfficeQA Pro evaluates AI models on professional knowledge-work questions and tasks drawn from real office workflows, including document analysis, spreadsheet reasoning, and information synthesis across business domains.
OfficeQA Pro is a text benchmark evaluating models on reasoning, general, and agents tasks. LLM Stats tracks 7 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.7.
Compare leaders on the best AI for reasoning, best AI for general and best AI for agents leaderboards.
Current leaders
Seed 2.1 Pro from ByteDance currently leads the OfficeQA Pro leaderboard with a score of 0.722 across 7 evaluated AI models.
FAQ
Common questions about the OfficeQA Pro benchmark and leaderboard.