The AI arena is free today

Open Superagent

SimpleQA

Paper

Progress Over Time

Interactive timeline showing model performance evolution on SimpleQA

State-of-the-art frontier
Open
Proprietary

SimpleQA Leaderboard

47 models
ContextCostLicense
1685B
2
3671B
4671B
5
6
71.0M$0.50 / $3.00
8
OpenAI
OpenAI
91.6T1.0M$1.60 / $3.20
10
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
235B
12
13
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B
141.0M$1.25 / $10.00
15
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
16
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B262K$0.10 / $0.60
17
OpenAI
OpenAI
18
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B
191.0M$0.25 / $1.50
20
21
OpenAI
OpenAI
128K$2.50 / $10.00
22
Moonshot AI
Moonshot AI
1.0T
23284B1.0M$0.14 / $0.28
241.0T
24
Moonshot AI
Moonshot AI
1.0T
26284B1.0M$0.10 / $0.20
27
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
31B
28
29
DeepSeek
DeepSeek
671B
30
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
31B
31675B
31675B
31675B262K$0.50 / $1.50
31675B
35
36456B
37456B
38
OpenAI
OpenAI
3924B
40
4124B
4227B
4312B
444B
45
Microsoft
Microsoft
15B
461B
4721B
Notice missing or incorrect data?

Sub-benchmarks

About this benchmark

What is SimpleQA?

SimpleQA is a factuality benchmark developed by OpenAI that measures the short-form factual accuracy of large language models. The benchmark contains 4,326 short, fact-seeking questions that are adversarially collected and designed to have single, indisputable answers. Questions cover diverse topics from science and technology to entertainment, and the benchmark also measures model calibration by evaluating whether models know what they know.

SimpleQA is a text benchmark evaluating models on reasoning, factuality, and general tasks. LLM Stats tracks 47 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 1.0.

Compare leaders on the best AI for reasoning, best AI for factuality and best AI for general leaderboards.

Current leaders

DeepSeek-V3.2-Exp from DeepSeek currently leads the SimpleQA leaderboard with a score of 0.971 across 47 evaluated AI models.

1DeepSeek-V3.2-ExpDeepSeek97.1%
2Grok 4 FastxAI95.0%
3DeepSeek-V3.1DeepSeek93.4%

Source paper

Title
Measuring short-form factuality in large language models
Authors
Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, and 4 others
Published
Abstract

We present SimpleQA, a benchmark that evaluates the ability of language models to answer short, fact-seeking questions. We prioritized two properties in designing this eval. First, SimpleQA is challenging, as it is adversarially collected against GPT-4 responses. Second, responses are easy to grade, because questions are created such that there exists only a single, indisputable answer. Each answer in SimpleQA is graded as either correct, incorrect, or not attempted. A model with ideal behavior would get as many questions correct as possible while not attempting the questions for which it is not confident it knows the correct answer. SimpleQA is a simple, targeted evaluation for whether models "know what they know," and our hope is that this benchmark will remain relevant for the next few generations of frontier models. SimpleQA can be found at https://github.com/openai/simple-evals.

FAQ

Common questions about the SimpleQA benchmark and leaderboard.

What is the SimpleQA benchmark?

SimpleQA is a factuality benchmark developed by OpenAI that measures the short-form factual accuracy of large language models. The benchmark contains 4,326 short, fact-seeking questions that are adversarially collected and designed to have single, indisputable answers. Questions cover diverse topics from science and technology to entertainment, and the benchmark also measures model calibration by evaluating whether models know what they know.

What is the SimpleQA leaderboard?

The SimpleQA leaderboard ranks 47 AI models based on their performance on this benchmark. Currently, DeepSeek-V3.2-Exp by DeepSeek leads with a score of 0.971. The average score across all models is 0.380.

What is the highest SimpleQA score?

The highest SimpleQA score is 0.971, achieved by DeepSeek-V3.2-Exp from DeepSeek.

How many models are evaluated on SimpleQA?

47 models have been evaluated on the SimpleQA benchmark, with 0 verified results and 47 self-reported results.

Where can I find the SimpleQA paper?

The SimpleQA paper is available at https://arxiv.org/abs/2411.04368. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does SimpleQA cover?

SimpleQA is categorized under reasoning, factuality, and general. The benchmark evaluates text models.

Are there variants of SimpleQA?

Yes. SimpleQA has 1 related variant: SimpleQA Verified.

What is the best open-source model on SimpleQA?

DeepSeek-V3.2-Exp by DeepSeek is the top-ranked open-source model on SimpleQA, with a score of 0.971 (rank #1).

How recent are the SimpleQA leaderboard results?

The SimpleQA leaderboard was last updated in August 2026 and currently includes 47 evaluated models.