Research & prototyping
MedGemma 4B IT
4/4 selected evidence sets represented. Best-supported starting point in this dataset; validate on your own images and task.
View model evidenceFor model research, MedGemma 4B IT has the strongest evidence coverage in LLM Stats, appearing on all 4 selected imaging benchmarks. For clinical use, choose an authorized device whose exact intended use matches the anatomy and finding—not a general chatbot or this model rank.
4 selected benchmarks
4 models represented
Hourly data refresh
No paid placement
Review boundary
Data methodology reviewed internally; the analysis has not been independently reviewed by a radiologist, medical physicist, or regulator.
Scope and disclosure
This page compares underlying model evidence for chest X-ray and radiology tasks. It does not rank clinical products, interpret patient images, or provide medical advice. LLM Stats accepts no payment for ranking position. Read the methodology or report a correction.
For research, compare the imaging benchmarks below. Hospitals need an authorized product for the intended finding. For your own X-ray, start with the radiology report and your clinician.
Research & prototyping
4/4 selected evidence sets represented. Best-supported starting point in this dataset; validate on your own images and task.
View model evidenceHospital or imaging center
Start with body part, projection, finding, user, workflow, and jurisdiction. Then verify authorization, external validation, PACS fit, monitoring, and human oversight.
Use the product checklistYour own X-ray
A consumer chatbot cannot establish a diagnosis or rule out an urgent finding. Protect identifying data and ask the ordering clinician or radiologist to interpret the image and report.
Understand the limitsLLM Stats shows each reported normalized result separately. CheXpert-CXR and MIMIC-CXR are chest-specific; VQA-RAD and SLAKE test question answering across mixed radiology images. Their metrics, samples, and tasks are not interchangeable.
Chest radiograph finding classification
48.1%
MedGemma reported score
Chest images paired with radiology reports
88.9%
MedGemma reported score
Clinically written questions about radiology images
49.9%
MedGemma reported score
Bilingual medical visual question answering
62.3%
MedGemma reported score
Scores are displayed as normalized values from each record. They must not be averaged or interpreted as probability that a clinical finding is correct. Current model coverage is sparse, and these selected records are marked unverified in LLM Stats at this refresh. Inspect the linked records and primary dataset sources before relying on them.
The available head-to-head
Useful for medical visual Q&A; not an X-ray-only comparison and not a clinical reader study.
“Analyzes X-rays” is not a usable requirement. A safe shortlist begins with the modality, anatomy, view, finding, intended user, action, and care setting.
ExamplesChest AP/PA, lateral, extremity, dental
Evidence neededExact image type, acquisition setting, prevalence
Common mismatchA chest model is treated as a universal X-ray reader
ExamplesTriage, detection, measurement, report draft
Evidence neededPer-finding sensitivity/specificity and threshold
Common mismatchA good report score is assumed to mean safe triage
ExamplesRadiologist, ED clinician, technologist, patient
Evidence neededWorkflow study with intended users and time pressure
Common mismatchResearch UI is deployed as clinical software
ExamplesAdult, pediatric, inpatient, outpatient
Evidence neededExternal and subgroup results with confidence intervals
Common mismatchOne-hospital performance is generalized everywhere
ExamplesDICOM, PACS/RIS, worklist, audit logs
Evidence neededLatency, failure handling, versioning, monitoring
Common mismatchThe model is compared without the finished product
A credible test compares unaided and AI-assisted reading on a locked, representative study set. Pre-register the endpoints and failure rules before looking at results.
Name the anatomy, projections, finding, user, action, and setting. One study should answer one intended-use question.
Use de-identified cases with reference standards, realistic prevalence, normal exams, mimics, devices, artifacts, and poor-quality acquisitions.
Sensitivity, specificity, AUROC where appropriate, false positives per image, calibration, time-to-read, and critical miss rate.
Randomize unaided versus AI-assisted reading, manage washout and order effects, and include readers with relevant experience.
Report site, equipment, age, sex, race or ethnicity where justified, projection, acuity, and image-quality strata with uncertainty.
Define unacceptable missed urgent findings, automation bias, unsafe prioritization, and system failures before calculating an average.
Model evidence is only one layer. Procurement should evaluate the marketed product, its precise claim, its version, and its behavior inside your workflow.
Exact modality, anatomy, finding, user, population, workflow, and contraindications.
Jurisdiction, submission or authorization record, cleared version, labeling, and whether your use falls inside the claim.
External sites, prospective or retrospective design, reference standard, prevalence, subgroups, confidence intervals, and conflicts.
DICOM and PHI flow, retention, subprocessors, residency, access, encryption, incident process, and training-data terms.
Alert burden, automation bias, explainability appropriate to the user, override, downtime, and escalation paths.
Version locking, change notices, drift monitoring, local quality metrics, rollback, audit logs, and accountable owners.
It cannot diagnose a person, clear a product, or prove better patient outcomes. The available model comparison is unusually sparse, so coverage is more defensible than a fabricated aggregate winner score.
Chest, dental, skeletal, mammography, fluoroscopy, and other radiographic tasks differ. Evidence on one does not transfer automatically to another.
Classification, report generation, and visual question answering use different samples and metrics. Their percentages are displayed separately for a reason.
A photograph or compressed export may omit DICOM information, views, priors, acquisition context, and clinical history. A chatbot response is not a radiology interpretation.
An image can contain identifiers in headers or pixels. Do not upload patient data until the complete service and workflow are approved for that use.
Direct answers, with the boundary between research evidence and clinical use kept intact.
For research models in the current LLM Stats evidence set, MedGemma 4B IT has the broadest coverage across the four selected chest X-ray and radiology evaluations. For clinical use, the best product is an authorized device validated for the exact anatomy, finding, population, user, and workflow.
A general chatbot output should not be used to diagnose, rule out, or manage a medical condition from an X-ray. Image quality, DICOM data, prior studies, history, and professional interpretation all matter. Use the official radiology report and discuss it with the ordering clinician or radiologist.
The current comparable public model evidence is too sparse to support one. Only SLAKE has several represented models, and it is a mixed-modality medical VQA test rather than an X-ray-only clinical study. Publishing ten confident winners would overstate the data.
It depends on the claim. Typical measures include per-finding sensitivity and specificity, AUROC where suitable, false positives per image, calibration, critical misses, reading time, subgroup performance, failure rate, and performance with versus without AI. Confidence intervals and external validation matter.
No. Authorization applies to the labeled device, version, intended use, users, and conditions described in its record. Verify that the proposed workflow fits that scope and still perform local implementation and monitoring work.