The AI arena is free today

Open Superagent

AI & LLM Benchmarks 2026

Explore live AI and LLM benchmark rankings across reasoning, coding, math, vision, agents, and tool use.

Compare composite indexes first, then open individual evaluations for score provenance, coverage, and methodology.

680 benchmarks55 capabilitiesData checked Sep 10, 2:47 PM UTC·refreshes every 30 min

AI model indexes

TrueSkill conservative ratings aggregated across benchmarks in each capability. Choose an index to see its current top ten.

Full overall leaderboard

Bar lengths show the spread within this top-ten field, not a zero-based scale. Higher conservative rating is better.

Broad capability across every indexed benchmark category.

Browse by task

Popular benchmark families

Methodology & limitations

We preserve benchmark-specific scores and distinguish reported from independently verified results where source data allows. Rankings can still move with prompt format, harness version, contamination, missing runs, and model updates; use several relevant tests rather than one universal score.

Scoring methodology