The AI arena is free today

Open Superagent
v3.1Updated September 2, 2026

LLM Stats Score methodology

A high-level overview of how we combine heterogeneous model benchmarks while accounting for uncertainty and incomplete coverage.

01

Overview

The LLM Stats Score combines comparable benchmark evidence into one conservative, uncertainty-aware view of model performance.

Missing results are not scored as failures. Models with less evidence retain greater uncertainty, keeping incomplete coverage distinct from observed performance.

02

Score construction

Benchmark scales do not carry the same meaning, so we compare relative performance within eligible benchmarks before combining evidence across the catalog.

Ties and metric direction are handled consistently while original benchmark values remain visible. The score favors broad, strong evidence and remains cautious when coverage is limited.

03

Evidence policy

Eligibility and missingness

Public benchmarks and approved community evaluations are considered when their results are sufficiently comparable. Missing results remain missing and are never turned into synthetic losses.

Provenance and precedence

Lab-reported values remain labeled and linked to their source. When LLM Stats independently verifies a result, the verified measurement is clearly identified.

04

Category indexes

Category indexes apply the same scoring principles to relevant benchmark subsets. The overall score is a separate view across the broader eligible evidence base.

Reasoning, Coding, Agent, Math, Long Context, and Vision provide independent views. A model can be highly rated in one category and remain unrated in another; absence means insufficient evidence, not a score of zero.

05

Limitations

Coverage

The score reflects the evidence available today. It cannot represent capabilities absent from the evaluation set and is not a task-specific deployment recommendation.

Correlated evidence

Related benchmarks can emphasize areas with denser public evaluation. We review the catalog as coverage evolves.

Reporting effects

Self-reported scores depend on evaluation choices. Provenance is visible, but labels cannot remove every source of variation.

Interpretation

Displayed uncertainty describes evidence inside this rating system, not every real-world task or future model behavior.

06

Further details

Need more information?

We share additional methodology details with researchers, customers, and partners when appropriate. Contact us to discuss evaluation design, score interpretation, provenance, or a specific use case.

Contact us for further details

Interpretation notes

We compare relative benchmark performance because score intervals do not carry consistent meaning across different evaluations.

Missing benchmarks do not count as losses. Lack of evaluation should reduce confidence, not be confused with observed failure.