The AI arena is free today

Open Superagent

GPQA

Progress Over Time

Interactive timeline showing model performance evolution on GPQA

State-of-the-art frontier
Open
Proprietary

GPQA Leaderboard

250 models
ContextCostLicense
11.1M$10.00 / $50.00
2
21.1M$5.00 / $30.00
41.0M$2.00 / $12.00
51.0M$5.00 / $25.00
6
OpenAI
OpenAI
1.1M$5.00 / $30.00
61.0M$5.00 / $25.00
8
Moonshot AI
Moonshot AI
2.8T1.0M$2.85 / $14.25
9
10500K$2.00 / $6.00
111.1M$2.00 / $12.00
12
OpenAI
OpenAI
1.0M$2.50 / $15.00
13
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
14
OpenAI
OpenAI
400K$1.75 / $14.00
14
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
16770B
161.1M$0.20 / $1.20
18
19
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B1.0M$0.15 / $0.47
19
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B
211.0M$5.00 / $25.00
22
Zhipu AI
Zhipu AI
753B1.0M$0.75 / $2.40
23763B1.0M$0.22 / $0.66
24
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
25
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
251.0M$0.50 / $3.00
25
Tencent
Tencent
295B262K$0.14 / $0.58
28
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
291.6T1.0M$1.30 / $2.60
301.0M$3.00 / $15.00
31
Thinking Machines Lab
Thinking Machines Lab
276B524K$0.30 / $1.20
31
33
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.40 / $3.00
34524K$0.30 / $1.20
35
ByteDance
ByteDance
256K$0.50 / $3.00
36
36
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B262K$0.45 / $3.00
38284B1.0M$0.09 / $0.18
38
38400K$1.25 / $10.00
38
OpenAI
OpenAI
400K$1.25 / $10.00
38
38
44400K$0.75 / $4.50
45
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.32 / $3.20
46
Moonshot AI
Moonshot AI
1.0T
47
48
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0T256K$1.20 / $6.00
48284B1.0M$0.09 / $0.18
50
150 of 250
1/5
Notice missing or incorrect data?
About this benchmark

What is GPQA?

A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Questions are Google-proof and extremely difficult, with PhD experts reaching 65% accuracy.

GPQA is a text benchmark evaluating models on physics, reasoning, general, biology, and chemistry tasks. LLM Stats tracks 250 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 1.0.

Compare leaders on the best AI for physics, best AI for reasoning, best AI for general, best AI for biology and best AI for chemistry leaderboards.

Current leaders

GPT-6 Astra from OpenAI currently leads the GPQA leaderboard with a score of 0.960 across 250 evaluated AI models.

1GPT-6 AstraOpenAI96.0%
2Claude Mythos PreviewAnthropic94.6%
2GPT-5.6 SolOpenAI94.6%
OSSHy4 preview#16 open-weight92.3%

Source paper

Title
GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Authors
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, and 4 others
Published
Abstract

We present GPQA, a challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. We ensure that the questions are high-quality and extremely difficult: experts who have or are pursuing PhDs in the corresponding domains reach 65% accuracy (74% when discounting clear mistakes the experts identified in retrospect), while highly skilled non-expert validators only reach 34% accuracy, despite spending on average over 30 minutes with unrestricted access to the web (i.e., the questions are "Google-proof"). The questions are also difficult for state-of-the-art AI systems, with our strongest GPT-4 based baseline achieving 39% accuracy. If we are to use future AI systems to help us answer very hard questions, for example, when developing new scientific knowledge, we need to develop scalable oversight methods that enable humans to supervise their outputs, which may be difficult even if the supervisors are themselves skilled and knowledgeable. The difficulty of GPQA both for skilled non-experts and frontier AI systems should enable realistic scalable oversight experiments, which we hope can help devise ways for human experts to reliably get truthful information from AI systems that surpass human capabilities.

FAQ

Common questions about the GPQA benchmark and leaderboard.

What is the GPQA benchmark?

A challenging dataset of 448 multiple-choice questions written by domain experts in biology, physics, and chemistry. Questions are Google-proof and extremely difficult, with PhD experts reaching 65% accuracy.

What is the GPQA leaderboard?

The GPQA leaderboard ranks 250 AI models based on their performance on this benchmark. Currently, GPT-6 Astra by OpenAI leads with a score of 0.960. The average score across all models is 0.686.

What is the highest GPQA score?

The highest GPQA score is 0.960, achieved by GPT-6 Astra from OpenAI.

How many models are evaluated on GPQA?

250 models have been evaluated on the GPQA benchmark, with 0 verified results and 247 self-reported results.

Where can I find the GPQA paper?

The GPQA paper is available at https://arxiv.org/abs/2311.12022. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does GPQA cover?

GPQA is categorized under physics, reasoning, general, biology, and chemistry. The benchmark evaluates text models.

What is the best open-source model on GPQA?

Hy4 preview by Tencent is the top-ranked open-source model on GPQA, with a score of 0.923 (rank #16).

Which model offers the best value on GPQA?

Among models scoring within 10% of the leader, DeepSeek-V4-Flash-Max from DeepSeek is the cheapest, at $0.09 per million input tokens with a score of 0.881.

How is GPQA scored?

GPQA is scored using accuracy, reported on a 0–1 scale. Lower is better only when explicitly noted; on this leaderboard, higher scores indicate better performance.

How recent are the GPQA leaderboard results?

The GPQA leaderboard was last updated in September 2026 and currently includes 250 evaluated models.