Overview
The LLM Stats Score combines comparable benchmark evidence into one conservative, uncertainty-aware view of model performance.
Missing results are not scored as failures. Models with less evidence retain greater uncertainty, keeping incomplete coverage distinct from observed performance.
Score construction
Benchmark scales do not carry the same meaning, so we compare relative performance within eligible benchmarks before combining evidence across the catalog.
Ties and metric direction are handled consistently while original benchmark values remain visible. The score favors broad, strong evidence and remains cautious when coverage is limited.
Evidence policy
Eligibility and missingness
Public benchmarks and approved community evaluations are considered when their results are sufficiently comparable. Missing results remain missing and are never turned into synthetic losses.
Provenance and precedence
Lab-reported values remain labeled and linked to their source. When LLM Stats independently verifies a result, the verified measurement is clearly identified.
Category indexes
Category indexes apply the same scoring principles to relevant benchmark subsets. The overall score is a separate view across the broader eligible evidence base.
Reasoning, Coding, Agent, Math, Long Context, and Vision provide independent views. A model can be highly rated in one category and remain unrated in another; absence means insufficient evidence, not a score of zero.
Limitations
Coverage
The score reflects the evidence available today. It cannot represent capabilities absent from the evaluation set and is not a task-specific deployment recommendation.
Correlated evidence
Related benchmarks can emphasize areas with denser public evaluation. We review the catalog as coverage evolves.
Reporting effects
Self-reported scores depend on evaluation choices. Provenance is visible, but labels cannot remove every source of variation.
Interpretation
Displayed uncertainty describes evidence inside this rating system, not every real-world task or future model behavior.
Further details
Need more information?
We share additional methodology details with researchers, customers, and partners when appropriate. Contact us to discuss evaluation design, score interpretation, provenance, or a specific use case.
Contact us for further detailsInterpretation notes
We compare relative benchmark performance because score intervals do not carry consistent meaning across different evaluations.
Missing benchmarks do not count as losses. Lack of evaluation should reduce confidence, not be confused with observed failure.