The AI arena is free today

Open Playground

AA-LCR

Progress Over Time

Interactive timeline showing model performance evolution on AA-LCR

State-of-the-art frontier
Open
Proprietary

AA-LCR Leaderboard

15 models
ContextCostLicense
1
Tencent
Tencent
295B
2
Mistral AI
Mistral AI
119B256K$0.15 / $0.60
3
Moonshot AI
Moonshot AI
1.0T
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B
7
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.30 / $2.40
8550B
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
10230B1.0M$0.30 / $1.20
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
12120B
13
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B
14
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2B
15
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
800M
Notice missing or incorrect data?
About this benchmark

What is AA-LCR?

Agent Arena Long Context Reasoning benchmark

AA-LCR is a text benchmark evaluating models on reasoning and long context tasks. LLM Stats tracks 15 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.7.

Compare leaders on the best AI for reasoning and best AI for long context leaderboards.

Current leaders

Hy3 from Tencent currently leads the AA-LCR leaderboard with a score of 0.734 across 15 evaluated AI models.

1Hy3Tencent73.4%
2Mistral Small 4Mistral AI71.2%
3Kimi K2.5Moonshot AI70.0%

FAQ

Common questions about the AA-LCR benchmark and leaderboard.

What is the AA-LCR benchmark?

Agent Arena Long Context Reasoning benchmark

What is the AA-LCR leaderboard?

The AA-LCR leaderboard ranks 15 AI models based on their performance on this benchmark. Currently, Hy3 by Tencent leads with a score of 0.734. The average score across all models is 0.586.

What is the highest AA-LCR score?

The highest AA-LCR score is 0.734, achieved by Hy3 from Tencent.

How many models are evaluated on AA-LCR?

15 models have been evaluated on the AA-LCR benchmark, with 0 verified results and 15 self-reported results.

What categories does AA-LCR cover?

AA-LCR is categorized under reasoning and long context. The benchmark evaluates text models.

What is the best open-source model on AA-LCR?

Hy3 by Tencent is the top-ranked open-source model on AA-LCR, with a score of 0.734 (rank #1).

Which model offers the best value on AA-LCR?

Among models scoring within 10% of the leader, Mistral Small 4 from Mistral AI is the cheapest, at $0.15 per million input tokens with a score of 0.712.

How recent are the AA-LCR leaderboard results?

The AA-LCR leaderboard was last updated in August 2026 and currently includes 15 evaluated models.