AI Leaderboard — Compare 300+ Top AI Models by Intelligence, Speed & Price
Independent rankings of GPT, Claude, Gemini, Llama, DeepSeek and 300+ AI models — composite LLM Stats Score, updated continuously from public benchmarks and live API metrics. See the full LLM Leaderboard for complete LLM rankings with advanced filters.
leads on reasoning
GPT-5.6 Sol
94.6%gpqa
wins at coding
Claude Opus 5
27arena
cheapest in the top 10
Grok 4.5
$2.00/M tok
fastest output
Gemini 3.7 Flash
595tok/s
longest context window
Grok-4 Fast Reasoning
2.0M tokenstokens
best open-weights
Kimi K3
93.5%gpqa
| License | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
1 | OpenAI | 57.4 | 56.8 | 50.6 | 44.4 | 2,134 | — | 1.1M | 85c/s | $7.78 | Proprietary |
2 | Anthropic | 56.2 | 55.3 | 42.7 | 41.4 | 2,668 | — | 1.0M | 58c/s | $7.22 | Proprietary |
3 | Anthropic | 56.1 | 53.5 | 48.8 | 42.8 | 2,003 | — | 1.0M | 96c/s | $14.44 | Proprietary |
4 | Anthropic | 56.0 | 56.8 | 46.6 | 38.2 | — | — | — | — | — | Proprietary |
5 | Moonshot AI | 54.9 | 53.9 | 45.9 | 41.6 | 1,816 | 2.8T | 1.0M | 124c/s | $4.33 | Open Source |
6 | Zhipu AI | 54.7 | 54.9 | 45.4 | 41.9 | — | 753B | 1.0M | — | $1.73 | Proprietary |
7 | DeepSeek | 54.5 | 52.0 | 44.3 | 40.8 | — | 1.6T | 1.0M | 199c/s | $0.48 | Open Source |
8 | Alibaba Cloud / Qwen Team | 53.2 | 52.0 | 42.1 | 40.0 | — | 2.4T | — | — | — | Open Source |
9 | OpenAI | 52.8 | 51.1 | 46.4 | 40.9 | 1,185 | — | 1.1M | 119c/s | $3.11 | Proprietary |
10 | Anthropic | 51.9 | 51.4 | 43.8 | 37.0 | 1,689 | — | 1.0M | 107c/s | $7.22 | Proprietary |
11 | Meta | 51.7 | 52.3 | 37.8 | 37.9 | 1,070 | — | 1.0M | 223c/s | $1.58 | Proprietary |
12 | Google | 51.1 | 49.9 | 38.6 | 35.9 | — | — | 1.0M | 595c/s | $1.08 | Proprietary |
13 | Anthropic | 49.6 | 49.0 | 40.0 | 33.9 | 1,588 | — | 1.0M | 59c/s | $2.89 | Proprietary |
14 | OpenAI | 49.3 | 48.6 | 41.1 | 34.7 | 2,124 | — | 1.1M | 46c/s | $7.78 | Proprietary |
15 | DeepSeek | 48.7 | 45.3 | 38.5 | 35.7 | — | — | 1.0M | — | $0.27 | Proprietary |
1-15 of 359
Recent
New Models
Announced in the last 15 days.
Index
Performance Index
Composite TrueSkill ratings across published benchmarks.
FAQ
Quick answers for choosing, comparing and interpreting today's leading AI models.