AlignBench
Progress Over Time
Interactive timeline showing model performance evolution on AlignBench
AlignBench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 73B | — | — | ||
| 2 | DeepSeek | 236B | — | — | ||
| 3 | Alibaba Cloud / Qwen Team | 8B | — | — | ||
| 4 | Alibaba Cloud / Qwen Team | 8B | — | — |
What is AlignBench?
AlignBench is a comprehensive multi-dimensional benchmark for evaluating Chinese alignment of Large Language Models. It contains 8 main categories: Fundamental Language Ability, Advanced Chinese Understanding, Open-ended Questions, Writing Ability, Logical Reasoning, Mathematics, Task-oriented Role Play, and Professional Knowledge. The benchmark includes 683 real-scenario rooted queries with human-verified references and uses a rule-calibrated multi-dimensional LLM-as-Judge approach with Chain-of-Thought for evaluation.
AlignBench is a text benchmark evaluating models on language, math, reasoning, roleplay, general, creativity, and writing tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.
Compare leaders on the best AI for language, best AI for math, best AI for reasoning, best AI for roleplay, best AI for general, best AI for creativity and best AI for writing leaderboards.
Current leaders
Qwen2.5 72B Instruct from Alibaba Cloud / Qwen Team currently leads the AlignBench leaderboard with a score of 0.816 across 4 evaluated AI models.
FAQ
Common questions about the AlignBench benchmark and leaderboard.