The AI arena is free today

Open Superagent

HLE-Verified

Progress Over Time

Interactive timeline showing model performance evolution on HLE-Verified

State-of-the-art frontier
Open
Proprietary

HLE-Verified Leaderboard

1 models
ContextCostLicense
11.0M$0.75 / $3.75
Notice missing or incorrect data?
About this benchmark

What is HLE-Verified?

HLE-Verified evaluates multidisciplinary expert reasoning on a verified subset of Humanity's Last Exam.

HLE-Verified is a text benchmark evaluating models on knowledge and reasoning tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.5.

Compare leaders on the best AI for knowledge and best AI for reasoning leaderboards.

Current leaders

Gemini 3.7 Flash from Google currently leads the HLE-Verified leaderboard with a score of 0.536 across 1 evaluated AI models.

1Gemini 3.7 FlashGoogle53.6%

FAQ

Common questions about the HLE-Verified benchmark and leaderboard.

What is the HLE-Verified benchmark?

HLE-Verified evaluates multidisciplinary expert reasoning on a verified subset of Humanity's Last Exam.

What is the HLE-Verified leaderboard?

The HLE-Verified leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Gemini 3.7 Flash by Google leads with a score of 0.536. The average score across all models is 0.536.

What is the highest HLE-Verified score?

The highest HLE-Verified score is 0.536, achieved by Gemini 3.7 Flash from Google.

How many models are evaluated on HLE-Verified?

1 models have been evaluated on the HLE-Verified benchmark, with 0 verified results and 1 self-reported results.

What categories does HLE-Verified cover?

HLE-Verified is categorized under knowledge and reasoning. The benchmark evaluates text models.

What's the difference between HLE-Verified and Humanity's Last Exam?

HLE-Verified is a variant of Humanity's Last Exam. See the Humanity's Last Exam leaderboard for the broader benchmark and per-model comparison.

Which model offers the best value on HLE-Verified?

Among models scoring within 10% of the leader, Gemini 3.7 Flash from Google is the cheapest, at $0.75 per million input tokens with a score of 0.536.

How recent are the HLE-Verified leaderboard results?

The HLE-Verified leaderboard was last updated in August 2026 and currently includes 1 evaluated models.