# Best AI for X-ray analysis evidence methodology

Retrieved: 2026-09-05T06:49:43.707Z

License: CC BY 4.0

Canonical analysis: https://llm-stats.com/best-ai-for-x-ray-analysis

## The question this evidence can answer

This dataset compares available underlying model results on four selected chest X-ray and radiology evaluations: chexpert-cxr, mimic-cxr, vqa-rad, slakevqa. It can identify which models have the broadest represented benchmark coverage and expose reported results within each evaluation.

It does not rank marketed clinical products, medical devices, hospitals, or radiologists. It cannot interpret a patient's image, establish a diagnosis, demonstrate clinical utility, prove improved outcomes, or determine regulatory status.

## Why there is no aggregate score

CheXpert-CXR, MIMIC-CXR, VQA-RAD, and SLAKE differ in task, image mix, samples, labels, questions, and metrics. LLM Stats therefore reports each result separately and ranks only evidence coverage. Averaging the percentages would create a number with no defensible clinical meaning.

## Reproduce the evidence view

1. Open https://llm-stats.com/best-ai-for-x-ray-analysis and record the retrieval time.
2. Download https://llm-stats.com/research/best-ai-for-x-ray-analysis/evidence.csv or https://llm-stats.com/research/best-ai-for-x-ray-analysis/evidence.json during the same hourly refresh window.
3. Confirm the selected benchmark IDs, model IDs, benchmark-level ranks, reported scores, normalized scores, and coverage counts against the server-rendered page and linked benchmark records.
4. Preserve missing data as missing. Do not infer a model's performance on an evaluation where no record is present.
5. Re-run after the data refreshes and document every changed record.

## Reader-study protocol

1. Lock one intended use: modality, anatomy, projection, target finding, user, population, setting, action, and product version.
2. Pre-register the primary endpoints, analysis plan, subgroups, sample size rationale, and severe-failure rules.
3. Assemble a de-identified, representative image set with an appropriate reference standard. Include normal exams, mimics, devices, artifacts, difficult cases, image-quality failures, and realistic prevalence.
4. Randomize unaided and AI-assisted reading, manage washout and order effects, and include readers with experience relevant to the intended use.
5. Measure sensitivity, specificity, AUROC when appropriate, false positives per image, calibration, critical misses, reading time, system failures, and automation-bias events. Report confidence intervals.
6. Analyze prespecified sites, devices, acquisition settings, demographic groups, projections, acuity levels, and image-quality strata where scientifically and ethically justified.
7. Preserve the locked dataset, case inclusion flow, reference decisions, reader assignments, raw outputs, exclusions, code, product settings, and failures.
8. Obtain the applicable clinical, privacy, security, legal, regulatory, and institutional review before deployment. Monitor the locked product version after launch and repeat evaluation after material change.

Download the reusable row-level template at https://llm-stats.com/research/best-ai-for-x-ray-analysis/reader-study-scorecard.csv.

## Limitations

- Current cross-model evidence is sparse; only one selected benchmark contains several represented models.
- VQA-RAD and SLAKE are mixed-radiology visual question-answering datasets, not X-ray-only clinical reader studies.
- Public benchmark performance may not transfer to a new institution, population, device, workflow, prevalence, or image-quality distribution.
- The analysis is maintained by LLM Stats and has not been independently reviewed by a radiologist, medical physicist, or regulator.
