The AI arena is free today

Open Superagent

AA-LCR

Progress Over Time

Interactive timeline showing model performance evolution on AA-LCR

State-of-the-art frontier
Open
Proprietary

AA-LCR Leaderboard

21 models
ContextCostLicense
130B131K$0.30 / $1.20
2
Tencent
Tencent
295B262K$0.14 / $0.58
3
Mistral AI
Mistral AI
119B256K$0.15 / $0.60
4524K$0.30 / $1.20
5
Moonshot AI
Moonshot AI
1.0T
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0T256K$1.20 / $6.00
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B262K$0.45 / $3.00
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B262K$0.29 / $2.40
10
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.26 / $2.60
11550B262K$0.50 / $2.20
12
InclusionAI
InclusionAI
124B131K$0.06 / $0.18
13
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B262K$0.10 / $0.15
14230B1.0M$0.30 / $1.20
15
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.14 / $1.00
16120B262K$0.09 / $0.40
17
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B
1830B262K$0.08 / $0.20
19
LG AI Research
LG AI Research
33B
20
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2B
21
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
800M
Notice missing or incorrect data?
About this benchmark

What is AA-LCR?

Agent Arena Long Context Reasoning benchmark

AA-LCR is a text benchmark evaluating models on long context and reasoning tasks. LLM Stats tracks 21 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.8.

Compare leaders on the best AI for long context and best AI for reasoning leaderboards.

Current leaders

Muse Glimmer-30B from Meta currently leads the AA-LCR leaderboard with a score of 0.800 across 21 evaluated AI models.

2Hy3Tencent73.4%
3Mistral Small 4Mistral AI71.2%

FAQ

Common questions about the AA-LCR benchmark and leaderboard.

What is the AA-LCR benchmark?

Agent Arena Long Context Reasoning benchmark

What is the AA-LCR leaderboard?

The AA-LCR leaderboard ranks 21 AI models based on their performance on this benchmark. Currently, Muse Glimmer-30B by Meta leads with a score of 0.800. The average score across all models is 0.603.

What is the highest AA-LCR score?

The highest AA-LCR score is 0.800, achieved by Muse Glimmer-30B from Meta.

How many models are evaluated on AA-LCR?

21 models have been evaluated on the AA-LCR benchmark, with 0 verified results and 21 self-reported results.

What categories does AA-LCR cover?

AA-LCR is categorized under long context and reasoning. The benchmark evaluates text models.

What is the best open-source model on AA-LCR?

Muse Glimmer-30B by Meta is the top-ranked open-source model on AA-LCR, with a score of 0.800 (rank #1).

Which model offers the best value on AA-LCR?

Among models scoring within 10% of the leader, Hy3 from Tencent is the cheapest, at $0.14 per million input tokens with a score of 0.734.

How recent are the AA-LCR leaderboard results?

The AA-LCR leaderboard was last updated in September 2026 and currently includes 21 evaluated models.