The AI arena is free today

Open Superagent

HumanEval-Average

Paper

Progress Over Time

Interactive timeline showing model performance evolution on HumanEval-Average

State-of-the-art frontier
Open
Proprietary

HumanEval-Average Leaderboard

1 models
ContextCostLicense
1
Mistral AI
Mistral AI
22B
Notice missing or incorrect data?
About this benchmark

What is HumanEval-Average?

A variant of the HumanEval benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics

HumanEval-Average is a text benchmark evaluating models on reasoning tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.6.

Compare leaders on the best AI for reasoning leaderboards.

Current leaders

Codestral-22B from Mistral AI currently leads the HumanEval-Average leaderboard with a score of 0.615 across 1 evaluated AI models.

1Codestral-22BMistral AI61.5%

Source paper

Title
Evaluating Large Language Models Trained on Code
Authors
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, and 54 others
Published
Abstract

We introduce Codex, a GPT language model fine-tuned on publicly available code from GitHub, and study its Python code-writing capabilities. A distinct production version of Codex powers GitHub Copilot. On HumanEval, a new evaluation set we release to measure functional correctness for synthesizing programs from docstrings, our model solves 28.8% of the problems, while GPT-3 solves 0% and GPT-J solves 11.4%. Furthermore, we find that repeated sampling from the model is a surprisingly effective strategy for producing working solutions to difficult prompts. Using this method, we solve 70.2% of our problems with 100 samples per problem. Careful investigation of our model reveals its limitations, including difficulty with docstrings describing long chains of operations and with binding operations to variables. Finally, we discuss the potential broader impacts of deploying powerful code generation technologies, covering safety, security, and economics.

FAQ

Common questions about the HumanEval-Average benchmark and leaderboard.

What is the HumanEval-Average benchmark?

A variant of the HumanEval benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics

What is the HumanEval-Average leaderboard?

The HumanEval-Average leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Codestral-22B by Mistral AI leads with a score of 0.615. The average score across all models is 0.615.

What is the highest HumanEval-Average score?

The highest HumanEval-Average score is 0.615, achieved by Codestral-22B from Mistral AI.

How many models are evaluated on HumanEval-Average?

1 models have been evaluated on the HumanEval-Average benchmark, with 0 verified results and 1 self-reported results.

Where can I find the HumanEval-Average paper?

The HumanEval-Average paper is available at https://arxiv.org/abs/2107.03374. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does HumanEval-Average cover?

HumanEval-Average is categorized under reasoning. The benchmark evaluates text models.

What is the best open-source model on HumanEval-Average?

Codestral-22B by Mistral AI is the top-ranked open-source model on HumanEval-Average, with a score of 0.615 (rank #1).

How recent are the HumanEval-Average leaderboard results?

The HumanEval-Average leaderboard was last updated in August 2026 and currently includes 1 evaluated models.