# Best AI Chatbots evidence methodology

Version: 2026-08-27  
Published: 2026-08-27  
Modified: 2026-08-27  
Index snapshot retrieved: 2026-08-27T23:42:00-04:00  
License: CC BY 4.0

## What is published

This asset combines two evidence classes that remain deliberately separate:

1. **Measured model evidence:** conservative LLM Stats TrueSkill index scores (μ − 3σ), ranks, field sizes, and games played from the live category-index backend.
2. **Documented product evidence:** feature, price, privacy, and data-control statements linked to official product documentation.

It does not publish consumer-product test scores, standardized-run excerpts, or claims that a listed consumer app routed to a particular model during a test. No comparable product-level run records exist in the repository.

## Representative-model rule

A representative model connects a product vendor to a current model record on LLM Stats so readers can inspect relevant model evidence. It is not proof of consumer-app routing. Perplexity is intentionally left unmapped because its consumer product routes across multiple selectable and automatic models. Each record states its mapping limitation.

## Index mapping

| LLM Stats index | What it can inform | What it cannot establish |
| --- | --- | --- |
| General | Broad model capability | Supporting model evidence only |
| Search | Research support—not product citation accuracy | Supporting model evidence only |
| Communication | Writing support | Supporting model evidence only |
| Code | Coding support | Supporting model evidence only |
| Long context | Large-file reasoning—not upload UX | Supporting model evidence only |
| Vision | Image understanding—not image generation UX | Supporting model evidence only |
| Tool calling | Tool-use support—not product recovery UX | Supporting model evidence only |

There is no comparable LLM Stats consumer-product index for cross-chat memory, voice interruption handling, native image-generation UX, product citation entailment, or visible failure recovery. Those fields remain documented feature assertions or not measured; they are never inferred from unrelated model benchmarks.

## Score interpretation

The conservative index score is μ − 3σ. It is not a percentage and scores from different index categories should not be averaged. Missing evidence is not zero. Low games-played counts imply greater uncertainty and remain visible in the asset. See the full LLM Stats Score methodology at https://llm-stats.com/methodology/llm-stats-score.

## Freshness

The index snapshot carries an explicit retrieval timestamp. Each product record carries a source-verification date and revision history. Page and Dataset dateModified values are generated from the latest verification record.
