AA-LCR
Progress Over Time
Interactive timeline showing model performance evolution on AA-LCR
State-of-the-art frontier
Open
Proprietary
AA-LCR Leaderboard
15 models
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Tencent | 295B | — | — | ||
| 2 | Mistral AI | 119B | 256K | $0.15 / $0.60 | ||
| 3 | Moonshot AI | 1.0T | — | — | ||
| 4 | Alibaba Cloud / Qwen Team | 397B | — | — | ||
| 5 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 6 | Alibaba Cloud / Qwen Team | 122B | — | — | ||
| 7 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.30 / $2.40 | ||
| 8 | 550B | — | — | |||
| 9 | Alibaba Cloud / Qwen Team | 9B | — | — | ||
| 10 | MiniMax | 230B | 1.0M | $0.30 / $1.20 | ||
| 11 | Alibaba Cloud / Qwen Team | 35B | — | — | ||
| 12 | 120B | — | — | |||
| 13 | Alibaba Cloud / Qwen Team | 4B | — | — | ||
| 14 | Alibaba Cloud / Qwen Team | 2B | — | — | ||
| 15 | Alibaba Cloud / Qwen Team | 800M | — | — |
Notice missing or incorrect data?
What is AA-LCR?
Agent Arena Long Context Reasoning benchmark
AA-LCR is a text benchmark evaluating models on reasoning and long context tasks. LLM Stats tracks 15 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.7.
Compare leaders on the best AI for reasoning and best AI for long context leaderboards.
Current leaders
Hy3 from Tencent currently leads the AA-LCR leaderboard with a score of 0.734 across 15 evaluated AI models.
FAQ
Common questions about the AA-LCR benchmark and leaderboard.