The AI arena is free today

Open Superagent

BrowseComp

Paper

Progress Over Time

Interactive timeline showing model performance evolution on BrowseComp

State-of-the-art frontier
Open
Proprietary

BrowseComp Leaderboard

67 models
ContextCostLicense
1
Shanghai AI Laboratory
Shanghai AI Laboratory
744B——
2—1.1M$10.00 / $50.00
3
Moonshot AI
Moonshot AI
2.8T1.0M$2.85 / $14.25
4
Anthropic
Anthropic
—1.0M$5.00 / $25.00
5—1.1M$5.00 / $30.00
6———
7—1.1M$2.00 / $12.00
8———
9
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
10
ByteDance
ByteDance
———
11—1.0M$2.00 / $12.00
12———
13—1.0M$2.00 / $10.00
14
OpenAI
OpenAI
—1.1M$5.00 / $30.00
15—1.0M$5.00 / $25.00
16
Tencent
Tencent
295B262K$0.14 / $0.58
17—1.0M$5.00 / $25.00
18
MiniMax
MiniMax
428B1.0M$0.28 / $1.10
191.6T1.0M$1.30 / $2.60
20—1.1M$0.20 / $1.20
21
OpenAI
OpenAI
—1.0M$2.50 / $15.00
22
Zhipu AI
Zhipu AI
754B203K$1.05 / $3.50
22—1.0M$5.00 / $25.00
24———
25
Thinking Machines Lab
Thinking Machines Lab
276B524K$0.30 / $1.20
26
ByteDance
ByteDance
—256K$0.50 / $3.00
27230B1.0M$0.30 / $1.20
28
Zhipu AI
Zhipu AI
744B200K$1.00 / $3.20
29
Moonshot AI
Moonshot AI
1.0T——
30—1.0M$3.00 / $15.00
31284B1.0M$0.09 / $0.18
32
InclusionAI
InclusionAI
124B131K$0.06 / $0.18
33
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B262K$0.45 / $3.00
33196B66K$0.10 / $0.40
35
ByteDance
ByteDance
—256K$0.25 / $2.00
36
OpenAI
OpenAI
—400K$1.75 / $14.00
37
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B262K$0.29 / $2.40
38230B1.0M$0.30 / $1.20
39
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.14 / $1.00
39
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.26 / $2.60
411.0T——
42309B——
43560B——
44
OpenAI
OpenAI
———
45
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0T256K$1.20 / $6.00
46284B1.0M$0.09 / $0.18
47
Zhipu AI
Zhipu AI
358B203K$0.40 / $1.75
48
OpenAI
OpenAI
———
49685B164K$0.26 / $0.38
49685B——
1–50 of 67
1/2
Notice missing or incorrect data?

Sub-benchmarks

BrowseComp Long Context 128k

A challenging benchmark for evaluating web browsing agents' ability to persistently navigate the internet and find hard-to-locate, entangled information. Comprises 1,266 questions requiring strategic reasoning, creative search, and interpretation of retrieved content, with short and easily verifiable answers.

text•Max 1

BrowseComp Long Context 256k

BrowseComp is a benchmark for measuring the ability of agents to browse the web, comprising 1,266 questions that require persistently navigating the internet in search of hard-to-find, entangled information. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers. The benchmark focuses on questions where answers are obscure, time-invariant, and well-supported by evidence scattered across the open web.

text•Max 1

BrowseComp-VL

BrowseComp-VL is the vision-language variant of BrowseComp, evaluating multimodal models on web browsing comprehension tasks that require processing visual web page content alongside text.

multimodal•Max 1

BrowseComp-zh

A high-difficulty benchmark purpose-built to comprehensively evaluate LLM agents on the Chinese web, consisting of 289 multi-hop questions spanning 11 diverse domains including Film & TV, Technology, Medicine, and History. Questions are reverse-engineered from short, objective, and easily verifiable answers, requiring sophisticated reasoning and information reconciliation beyond basic retrieval. The benchmark addresses linguistic, infrastructural, and censorship-related complexities in Chinese web environments.

text•Max 1
About this benchmark

What is BrowseComp?

BrowseComp is a benchmark comprising 1,266 questions that challenge AI agents to persistently navigate the internet in search of hard-to-find, entangled information. The benchmark measures agents' ability to exercise persistence in information gathering, demonstrate creativity in web navigation, and find concise, verifiable answers. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers.

BrowseComp is a text benchmark evaluating models on reasoning, search, and agents tasks. LLM Stats tracks 67 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.9.

Compare leaders on the best AI for reasoning, best AI for search and best AI for agents leaderboards.

Current leaders

Atria Dawn Preview from Shanghai AI Laboratory currently leads the BrowseComp leaderboard with a score of 0.925 across 67 evaluated AI models.

1Atria Dawn PreviewShanghai AI Laboratory92.5%
2GPT-6 AstraOpenAI91.5%
3Kimi K3Moonshot AI91.2%

Source paper

Title
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
Authors
Jason Wei, Zhiqing Sun, Spencer Papay, Scott McKinney, and 6 others
Published
Abstract

We present BrowseComp, a simple yet challenging benchmark for measuring the ability for agents to browse the web. BrowseComp comprises 1,266 questions that require persistently navigating the internet in search of hard-to-find, entangled information. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers. BrowseComp for browsing agents can be seen as analogous to how programming competitions are an incomplete but useful benchmark for coding agents. While BrowseComp sidesteps challenges of a true user query distribution, like generating long answers or resolving ambiguity, it measures the important core capability of exercising persistence and creativity in finding information. BrowseComp can be found at https://github.com/openai/simple-evals.

FAQ

Common questions about the BrowseComp benchmark and leaderboard.

What is the BrowseComp benchmark?

BrowseComp is a benchmark comprising 1,266 questions that challenge AI agents to persistently navigate the internet in search of hard-to-find, entangled information. The benchmark measures agents' ability to exercise persistence in information gathering, demonstrate creativity in web navigation, and find concise, verifiable answers. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers.

What is the BrowseComp leaderboard?

The BrowseComp leaderboard ranks 67 AI models based on their performance on this benchmark. Currently, Atria Dawn Preview by Shanghai AI Laboratory leads with a score of 0.925. The average score across all models is 0.652.

What is the highest BrowseComp score?

The highest BrowseComp score is 0.925, achieved by Atria Dawn Preview from Shanghai AI Laboratory.

How many models are evaluated on BrowseComp?

67 models have been evaluated on the BrowseComp benchmark, with 0 verified results and 67 self-reported results.

Where can I find the BrowseComp paper?

The BrowseComp paper is available at https://arxiv.org/abs/2504.12516. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does BrowseComp cover?

BrowseComp is categorized under reasoning, search, and agents. The benchmark evaluates text models.

Are there variants of BrowseComp?

What is the best open-source model on BrowseComp?

Atria Dawn Preview by Shanghai AI Laboratory is the top-ranked open-source model on BrowseComp, with a score of 0.925 (rank #1).

Which model offers the best value on BrowseComp?

Among models scoring within 10% of the leader, Hy3 from Tencent is the cheapest, at $0.14 per million input tokens with a score of 0.842.

How recent are the BrowseComp leaderboard results?

The BrowseComp leaderboard was last updated in October 2026 and currently includes 67 evaluated models.