The AI arena is free today

Open Superagent

AlignBench

Paper

Progress Over Time

Interactive timeline showing model performance evolution on AlignBench

State-of-the-art frontier
Open
Proprietary

AlignBench Leaderboard

4 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
73B
2236B
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
8B
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
8B
Notice missing or incorrect data?
About this benchmark

What is AlignBench?

AlignBench is a comprehensive multi-dimensional benchmark for evaluating Chinese alignment of Large Language Models. It contains 8 main categories: Fundamental Language Ability, Advanced Chinese Understanding, Open-ended Questions, Writing Ability, Logical Reasoning, Mathematics, Task-oriented Role Play, and Professional Knowledge. The benchmark includes 683 real-scenario rooted queries with human-verified references and uses a rule-calibrated multi-dimensional LLM-as-Judge approach with Chain-of-Thought for evaluation.

AlignBench is a text benchmark evaluating models on language, math, reasoning, roleplay, general, creativity, and writing tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.

Compare leaders on the best AI for language, best AI for math, best AI for reasoning, best AI for roleplay, best AI for general, best AI for creativity and best AI for writing leaderboards.

Current leaders

Qwen2.5 72B Instruct from Alibaba Cloud / Qwen Team currently leads the AlignBench leaderboard with a score of 0.816 across 4 evaluated AI models.

1Qwen2.5 72B InstructAlibaba Cloud / Qwen Team81.6%
2DeepSeek-V2.5DeepSeek80.4%
3Qwen2.5 7B InstructAlibaba Cloud / Qwen Team73.3%

FAQ

Common questions about the AlignBench benchmark and leaderboard.

What is the AlignBench benchmark?

AlignBench is a comprehensive multi-dimensional benchmark for evaluating Chinese alignment of Large Language Models. It contains 8 main categories: Fundamental Language Ability, Advanced Chinese Understanding, Open-ended Questions, Writing Ability, Logical Reasoning, Mathematics, Task-oriented Role Play, and Professional Knowledge. The benchmark includes 683 real-scenario rooted queries with human-verified references and uses a rule-calibrated multi-dimensional LLM-as-Judge approach with Chain-of-Thought for evaluation.

What is the AlignBench leaderboard?

The AlignBench leaderboard ranks 4 AI models based on their performance on this benchmark. Currently, Qwen2.5 72B Instruct by Alibaba Cloud / Qwen Team leads with a score of 0.816. The average score across all models is 0.769.

What is the highest AlignBench score?

The highest AlignBench score is 0.816, achieved by Qwen2.5 72B Instruct from Alibaba Cloud / Qwen Team.

How many models are evaluated on AlignBench?

4 models have been evaluated on the AlignBench benchmark, with 0 verified results and 4 self-reported results.

Where can I find the AlignBench paper?

The AlignBench paper is available at https://arxiv.org/abs/2311.18743. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does AlignBench cover?

AlignBench is categorized under language, math, reasoning, roleplay, general, creativity, and writing. The benchmark evaluates text models with multilingual support.

What is the best open-source model on AlignBench?

Qwen2.5 72B Instruct by Alibaba Cloud / Qwen Team is the top-ranked open-source model on AlignBench, with a score of 0.816 (rank #1).

How recent are the AlignBench leaderboard results?

The AlignBench leaderboard was last updated in August 2026 and currently includes 4 evaluated models.