Open LLM Leaderboard
Ranking the best open LLMs by performance, price, and speed
Open LLM Leaderboard highlights
Independent ranking of open-weight large language models — Llama, Qwen, GLM, DeepSeek, Mistral, Kimi and more — by coding-arena score, GPQA Diamond, throughput, latency, and per-token pricing. Updated continuously from provider APIs and verified benchmarks. See the LLM Stats Score methodology for how rankings are computed.
- Best for coding (Arena): DeepSeek-V4-Flash-0731 (24.1 arena score)
- Best on GPQA Diamond: Kimi K3 (93.5%)
- Best on AIME 2025: Kimi K2-Thinking-0905 (100.0%)
- Best on SWE-Bench Verified: DeepSeek-V4-Pro-Max (80.6%)
- Highest throughput: DeepSeek-V4-Flash-0731 (180 tok/s)
- Lowest latency: Mistral Small 4 (1413.90 s)
- Cheapest input: Mistral NeMo Instruct ($0.02 / 1M tok)
- Largest context window: Kimi K3 (1.0M tokens)
FAQ
Common questions about the open llm leaderboard