AI & LLM Benchmarks 2026
Explore live AI and LLM benchmark rankings across reasoning, coding, math, vision, agents, and tool use.
Compare composite indexes first, then open individual evaluations for score provenance, coverage, and methodology.
AI model indexes
TrueSkill conservative ratings aggregated across benchmarks in each capability. Choose an index to see its current top ten.
Bar lengths show the spread within this top-ten field, not a zero-based scale. Higher conservative rating is better.
Broad capability across every indexed benchmark category.
Browse by task
Popular benchmark families
Methodology & limitations
We preserve benchmark-specific scores and distinguish reported from independently verified results where source data allows. Rankings can still move with prompt format, harness version, contamination, missing runs, and model updates; use several relevant tests rather than one universal score.
Scoring methodology