The AI arena is free today

Open Superagent

Best by task

Best for Reasoning
GPT-6 Astra57.9 indexHighest reasoning index
Best Open-Weight
Kimi K393.5% GPQATop open-weight model

LLM Leaderboard 2026

Compare 300+ AI models by the LLM Stats Score — intelligence, speed and price, updated continuously from public benchmarks and live API metrics.

LLM Leaderboard highlights

The LLM Leaderboard ranks 300+ AI models by intelligence, output speed, latency and per-token pricing, aggregated into the LLM Stats Score. Updated continuously from provider APIs and verified benchmarks. See the LLM Stats Score methodology for how rankings are computed.

FAQ

Common questions about the llm leaderboard

What is the best LLM right now?

Based on coding-arena performance — the most discriminating signal at the frontier — Claude Opus 5 currently leads. For knowledge-heavy reasoning (GPQA Diamond), GPT-6 Astra scores highest. Choose by axis rather than a single ranking — see the highlights above for per-metric leaders.

How does the LLM Leaderboard rank models?

Models are sorted by coding-arena score (when available), then by GPQA Diamond. Each row aggregates verified benchmark results, provider-reported pricing, and live performance metrics (output throughput and time-to-first-token) sampled across the major API providers. See the LLM Stats Score methodology for a high-level overview and its limitations.

How many models are tracked?

This leaderboard tracks 390 canonical models across every major lab and provider. New releases typically appear within hours.

Where does pricing data come from?

Per-model input/output pricing is pulled from each provider's public API price list and verified against billing samples from the LLM Stats proxy. When a model is hosted by multiple providers, the cheapest available rate is shown by default.

How is performance measured?

Output throughput (tokens/second) and time-to-first-token are collected consistently across supported provider APIs and summarized over a recent observation window. Per-model splits live on each model detail page.

How often does the data update?

Pricing, model metadata, and live performance metrics are reviewed on a recurring basis. Benchmark scores update when new evidence is verified and published.