ScienceQA
Progress Over Time
Interactive timeline showing model performance evolution on ScienceQA
ScienceQA Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Microsoft | 4B | — | — |
What is ScienceQA?
ScienceQA is the first large-scale multimodal science question answering benchmark with 21,208 multiple-choice questions covering 3 subjects (natural science, language science, social science), 26 topics, 127 categories, and 379 skills. The benchmark includes both text and image modalities, featuring detailed explanations and Chain-of-Thought reasoning to diagnose multi-hop reasoning ability.
ScienceQA is a multimodal benchmark evaluating models on math, multimodal, reasoning, and vision tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 0.9.
Compare leaders on the best AI for math, best AI for multimodal, best AI for reasoning and best AI for vision leaderboards.
Current leaders
Phi-3.5-vision-instruct from Microsoft currently leads the ScienceQA leaderboard with a score of 0.913 across 1 evaluated AI models.
FAQ
Common questions about the ScienceQA benchmark and leaderboard.