The AI arena is free today

Open Superagent

Best AI for X-ray analysis

For model research, MedGemma 4B IT has the strongest evidence coverage in LLM Stats, appearing on all 4 selected imaging benchmarks. For clinical use, choose an authorized device whose exact intended use matches the anatomy and finding—not a general chatbot or this model rank.

4 selected benchmarks

4 models represented

Hourly data refresh

No paid placement

Evidence viewerIllustration · not a patient image
Stylized chest X-ray evidence mapA non-diagnostic illustration connecting a chest radiograph silhouette to four benchmark evidence labels.evidence regionCHEXPERT-CXRMIMIC-CXRVQA-RADSLAKE
4 evidence sets
Benchmark coverage, not a diagnosis overlayMetrics shown separately · never averaged
Reviewed by Jonathan ChavezSources & review

Editor

Co-Founder, LLM Stats · model evaluation and benchmark design

Review boundary

Data methodology reviewed internally; the analysis has not been independently reviewed by a radiologist, medical physicist, or regulator.

Scope and disclosure

This page compares underlying model evidence for chest X-ray and radiology tasks. It does not rank clinical products, interpret patient images, or provide medical advice. LLM Stats accepts no payment for ranking position. Read the methodology or report a correction.

X-ray AI for researchers, hospitals, and patients

For research, compare the imaging benchmarks below. Hospitals need an authorized product for the intended finding. For your own X-ray, start with the radiology report and your clinician.

01

Research & prototyping

MedGemma 4B IT

4/4 selected evidence sets represented. Best-supported starting point in this dataset; validate on your own images and task.

View model evidence
02

Hospital or imaging center

Match an authorized product

Start with body part, projection, finding, user, workflow, and jurisdiction. Then verify authorization, external validation, PACS fit, monitoring, and human oversight.

Use the product checklist
03

Your own X-ray

Use the report and a clinician

A consumer chatbot cannot establish a diagnosis or rule out an urgent finding. Protect identifying data and ask the ordering clinician or radiologist to interpret the image and report.

Understand the limits

Compare X-ray models by benchmark

LLM Stats shows each reported normalized result separately. CheXpert-CXR and MIMIC-CXR are chest-specific; VQA-RAD and SLAKE test question answering across mixed radiology images. Their metrics, samples, and tasks are not interchangeable.

CheXpert-CXRChest X-ray

Chest radiograph finding classification

48.1%

MedGemma reported score

MIMIC-CXRChest X-ray

Chest images paired with radiology reports

88.9%

MedGemma reported score

VQA-RADMixed radiology

Clinically written questions about radiology images

49.9%

MedGemma reported score

SLAKEMixed modalities

Bilingual medical visual question answering

62.3%

MedGemma reported score

Scores are displayed as normalized values from each record. They must not be averaged or interpreted as probability that a clinical finding is correct. Current model coverage is sparse, and these selected records are marked unverified in LLM Stats at this refresh. Inspect the linked records and primary dataset sources before relying on them.

The available head-to-head

SLAKE mixed-modality ranking

Useful for medical visual Q&A; not an X-ray-only comparison and not a clinical reader study.

Choose AI for the anatomy and finding

“Analyzes X-rays” is not a usable requirement. A safe shortlist begins with the modality, anatomy, view, finding, intended user, action, and care setting.

Modality & anatomy

ExamplesChest AP/PA, lateral, extremity, dental

Evidence neededExact image type, acquisition setting, prevalence

Common mismatchA chest model is treated as a universal X-ray reader

Finding & action

ExamplesTriage, detection, measurement, report draft

Evidence neededPer-finding sensitivity/specificity and threshold

Common mismatchA good report score is assumed to mean safe triage

User & setting

ExamplesRadiologist, ED clinician, technologist, patient

Evidence neededWorkflow study with intended users and time pressure

Common mismatchResearch UI is deployed as clinical software

Population

ExamplesAdult, pediatric, inpatient, outpatient

Evidence neededExternal and subgroup results with confidence intervals

Common mismatchOne-hospital performance is generalized everywhere

Integration

ExamplesDICOM, PACS/RIS, worklist, audit logs

Evidence neededLatency, failure handling, versioning, monitoring

Common mismatchThe model is compared without the finished product

Compare radiologists’ results with and without AI

A credible test compares unaided and AI-assisted reading on a locked, representative study set. Pre-register the endpoints and failure rules before looking at results.

01

Lock the claim

Name the anatomy, projections, finding, user, action, and setting. One study should answer one intended-use question.

02

Build the set

Use de-identified cases with reference standards, realistic prevalence, normal exams, mimics, devices, artifacts, and poor-quality acquisitions.

03

Pre-register endpoints

Sensitivity, specificity, AUROC where appropriate, false positives per image, calibration, time-to-read, and critical miss rate.

04

Compare reading modes

Randomize unaided versus AI-assisted reading, manage washout and order effects, and include readers with relevant experience.

05

Inspect subgroups

Report site, equipment, age, sex, race or ethnicity where justified, projection, acuity, and image-quality strata with uncertainty.

06

Gate severe failures

Define unacceptable missed urgent findings, automation bias, unsafe prioritization, and system failures before calculating an average.

Six checks before a clinical pilot

Model evidence is only one layer. Procurement should evaluate the marketed product, its precise claim, its version, and its behavior inside your workflow.

  1. 1

    Intended use

    Exact modality, anatomy, finding, user, population, workflow, and contraindications.

  2. 2

    Regulatory status

    Jurisdiction, submission or authorization record, cleared version, labeling, and whether your use falls inside the claim.

  3. 3

    Clinical evidence

    External sites, prospective or retrospective design, reference standard, prevalence, subgroups, confidence intervals, and conflicts.

  4. 4

    Privacy & security

    DICOM and PHI flow, retention, subprocessors, residency, access, encryption, incident process, and training-data terms.

  5. 5

    Human factors

    Alert burden, automation bias, explainability appropriate to the user, override, downtime, and escalation paths.

  6. 6

    Lifecycle controls

    Version locking, change notices, drift monitoring, local quality metrics, rollback, audit logs, and accountable owners.

What this page cannot tell from a benchmark

It cannot diagnose a person, clear a product, or prove better patient outcomes. The available model comparison is unusually sparse, so coverage is more defensible than a fabricated aggregate winner score.

No universal X-ray model+

Chest, dental, skeletal, mammography, fluoroscopy, and other radiographic tasks differ. Evidence on one does not transfer automatically to another.

No score equivalence+

Classification, report generation, and visual question answering use different samples and metrics. Their percentages are displayed separately for a reason.

No consumer diagnosis+

A photograph or compressed export may omit DICOM information, views, priors, acquisition context, and clinical history. A chatbot response is not a radiology interpretation.

No privacy presumption+

An image can contain identifiers in headers or pixels. Do not upload patient data until the complete service and workflow are approved for that use.

X-ray AI: accuracy, clinical use, and privacy

Direct answers, with the boundary between research evidence and clinical use kept intact.

What is the best AI for X-ray analysis?+

For research models in the current LLM Stats evidence set, MedGemma 4B IT has the broadest coverage across the four selected chest X-ray and radiology evaluations. For clinical use, the best product is an authorized device validated for the exact anatomy, finding, population, user, and workflow.

Can ChatGPT or another general chatbot read my X-ray?+

A general chatbot output should not be used to diagnose, rule out, or manage a medical condition from an X-ray. Image quality, DICOM data, prior studies, history, and professional interpretation all matter. Use the official radiology report and discuss it with the ordering clinician or radiologist.

Why is there no top-10 X-ray AI list here?+

The current comparable public model evidence is too sparse to support one. Only SLAKE has several represented models, and it is a mixed-modality medical VQA test rather than an X-ray-only clinical study. Publishing ten confident winners would overstate the data.

What metrics matter for clinical X-ray AI?+

It depends on the claim. Typical measures include per-finding sensitivity and specificity, AUROC where suitable, false positives per image, calibration, critical misses, reading time, subgroup performance, failure rate, and performance with versus without AI. Confidence intervals and external validation matter.

Does FDA authorization mean an AI works for every X-ray?+

No. Authorization applies to the labeled device, version, intended use, users, and conditions described in its record. Verify that the proposed workflow fits that scope and still perform local implementation and monitoring work.

Compare medical models and imaging benchmarks