How we measure AI.
A high-level overview of how LLM Stats collects, evaluates, and combines evidence about AI models.
- Updated
- September 2026
- Version
- v3.1
- Maintained at
- llm-stats.com
1Evidence.
LLM Stats combines independently evaluated results with clearly labelled external evidence. We prioritize consistent comparisons, source provenance, and version history so readers can understand what each result represents.
Evidence is reviewed before publication and may be updated when a model, benchmark, or source materially changes.
2Arenas.
Arenas capture human preferences through blind comparisons. The experience is designed to reduce the influence of model identity and presentation order while supporting a broad range of tasks and modalities.
Quality controls help identify unreliable activity and preserve the usefulness of aggregate preference signals.
3Rating approach.
Relative performance is combined with an uncertainty-aware statistical approach. Stronger and more consistent evidence has more influence, while limited or conflicting evidence is treated more cautiously.
Missing observations are not treated as failures. Ratings are intended for comparison within their stated context, not as a universal measure of model quality.
4Coverage.
The catalogue covers major capability areas including coding, reasoning, multimodal understanding, long-context work, and professional domains. Coverage evolves as evaluation practices and available models change.
Results from different domains remain distinct where combining them would hide meaningful differences in model behavior.
5Limitations.
Benchmarks and preference data are imperfect proxies for real deployments. Results can be affected by task selection, model updates, provider behavior, evaluation design, and the availability of comparable evidence.
Rankings should be considered alongside cost, latency, safety, reliability, and requirements specific to the intended use case.
6Further details.
We keep implementation parameters, operational controls, and other sensitive methodology details out of this public overview. Researchers, partners, and customers can contact us for further details.
7Acknowledgements.
Supported by
Y Combinator (S25), with angel investment from Ivan Burazin (founder, Daytona), Thomas Wolf (Hugging Face), researchers at Harvard Medical, and employees and executives at Google and Datadog.
Cited in
OpenAI announcements, NBC News, TechCrunch, The Lancet, and Seeking Alpha.
Maintainers
Maintained by Jonathan Chávez and Sebastian Crossa, co-founders, LLM Stats.
Cite this methodology
Chávez, J., & Crossa, S. (2026). How we measure AI: Methodology, v3.1. LLM Stats. https://llm-stats.com/research