Best by task
- Best for Reasoning
- GPT-6 Astra57.9 indexHighest reasoning index →
- Best for Coding
- Claude Opus 527 arenaTop coding arena score →
- Fastest LLM
- Gemini 3.5 Flash396 c/sHighest P95 output throughput →
- Cheapest Frontier
- Gemma 4 E4B$0.02/1M inLowest price at frontier quality →
- Largest Context
- Grok-4.1 Fast Non-Reasoning2M tokensBiggest context window →
- Best Open-Weight
- Kimi K393.5% GPQATop open-weight model →
LLM Leaderboard 2026
Compare 300+ AI models by the LLM Stats Score — intelligence, speed and price, updated continuously from public benchmarks and live API metrics.
LLM Leaderboard highlights
The LLM Leaderboard ranks 300+ AI models by intelligence, output speed, latency and per-token pricing, aggregated into the LLM Stats Score. Updated continuously from provider APIs and verified benchmarks. See the LLM Stats Score methodology for how rankings are computed.
- Best for coding (Arena): Claude Opus 5 (26.7 arena score)
- Best on GPQA Diamond: GPT-6 Astra (96.0%)
- Best on AIME 2025: GPT-5.2 Pro (100.0%)
- Best on SWE-Bench Verified: Claude Fable 5 (95.0%)
- Highest throughput: Gemini 3.5 Flash (396 tok/s)
- Lowest latency: Claude Haiku 4.5 (1092.50 s)
- Cheapest input: Mistral NeMo Instruct ($0.02 / 1M tok)
- Largest context window: Grok-4.1 Fast Non-Reasoning (2.0M tokens)
FAQ
Common questions about the llm leaderboard