GPQA

Progress Over Time

Interactive timeline showing model performance evolution on GPQA

State-of-the-art frontier
Open
Proprietary

GPQA Leaderboard

232 models
ContextCostLicense
11.1M$5.00 / $30.00
1
31.0M$2.50 / $15.00
41.0M$5.00 / $25.00
51.0M$5.00 / $25.00
5
OpenAI
OpenAI
1.1M$5.00 / $30.00
7
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
8
9500K$2.00 / $6.00
101.1M$2.50 / $15.00
11
OpenAI
OpenAI
1.0M$2.50 / $15.00
12
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
12
OpenAI
OpenAI
400K$1.75 / $14.00
141.1M$1.00 / $6.00
15
161.0M$5.00 / $25.00
17
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
18
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
19
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
19
Tencent
Tencent
295B
191.0M$0.50 / $3.00
22
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.32 / $1.28
231.6T1.0M$1.60 / $3.20
24200K$3.00 / $15.00
25
26
ByteDance
ByteDance
256K$0.50 / $3.00
27
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B
27
29
OpenAI
OpenAI
400K$1.25 / $10.00
29400K$1.25 / $10.00
29
29
29
29284B1.0M$0.10 / $0.20
35400K$0.75 / $4.50
36
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.60 / $3.60
37
Moonshot AI
Moonshot AI
1.0T
38
39
40
40550B
421.0M$0.25 / $1.50
43
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B
44
45
Zhipu AI
Zhipu AI
754B200K$1.40 / $4.40
46
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
47
OpenAI
OpenAI
47
Zhipu AI
Zhipu AI
358B
47
50400K$5.00 / $30.00
150 of 232
1/5
Notice missing or incorrect data?
About this benchmark

What is GPQA?

A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Questions are Google-proof and extremely difficult, with PhD experts reaching 65% accuracy.

GPQA is a text benchmark evaluating models on physics, reasoning, general, biology, and chemistry tasks. LLM Stats tracks 232 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.9.

Compare leaders on the best AI for physics, best AI for reasoning, best AI for general, best AI for biology and best AI for chemistry leaderboards.

Current leaders

GPT-5.6 Sol from OpenAI currently leads the GPQA leaderboard with a score of 0.946 across 232 evaluated AI models.

1GPT-5.6 SolOpenAI94.6%
1Claude Mythos PreviewAnthropic94.6%
3Gemini 3.1 ProGoogle94.3%
OSSKimi K3#7 open-weight93.5%

Source paper

Title
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Authors
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, and 4 others
Published
Abstract

We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect), while highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web (i.e., the questions are "Google-proof"). The questions are also difficult for state-of-the-art AI systems, with our strongest GPT-4 based baseline achieving 39% accuracy. If we are to use future AI systems to help us answer very hard questions, for example, when developing new scientific knowledge, we need to develop scalable oversight methods that enable humans to supervise their outputs, which may be difficult even if the supervisors are themselves skilled and knowledgeable. The difficulty of GPQA both for skilled non-experts and frontier AI systems should enable realistic scalable oversight experiments, which we hope can help devise ways for human experts to reliably get truthful information from AI systems that surpass human capabilities.

FAQ

Common questions about the GPQA benchmark and leaderboard.

What is the GPQA benchmark?

A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Questions are Google-proof and extremely difficult, with PhD experts reaching 65% accuracy.

What is the GPQA leaderboard?

The GPQA leaderboard ranks 232 AI models based on their performance on this benchmark. Currently, GPT-5.6 Sol by OpenAI leads with a score of 0.946. The average score across all models is 0.674.

What is the highest GPQA score?

The highest GPQA score is 0.946, achieved by GPT-5.6 Sol from OpenAI.

How many models are evaluated on GPQA?

232 models have been evaluated on the GPQA benchmark, with 0 verified results and 229 self-reported results.

Where can I find the GPQA paper?

The GPQA paper is available at https://arxiv.org/abs/2311.12022. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does GPQA cover?

GPQA is categorized under physics, reasoning, general, biology, and chemistry. The benchmark evaluates text models.

What is the best open-source model on GPQA?

Kimi K3 by Moonshot AI is the top-ranked open-source model on GPQA, with a score of 0.935 (rank #7).

Which model offers the best value on GPQA?

Among models scoring within 10% of the leader, DeepSeek-V4-Flash-Max from DeepSeek is the cheapest, at $0.10 per million input tokens with a score of 0.881.

How is GPQA scored?

GPQA is scored using accuracy, reported on a 0–1 scale. Lower is better only when explicitly noted; on this leaderboard, higher scores indicate better performance.

How recent are the GPQA leaderboard results?

The GPQA leaderboard was last updated in July 2026 and currently includes 232 evaluated models.