HumanEval-Average
Progress Over Time
Interactive timeline showing model performance evolution on HumanEval-Average
HumanEval-Average Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Mistral AI | 22B | — | — |
What is HumanEval-Average?
A variant of the HumanEval benchmark that measures functional correctness for synthesizing programs from docstrings, consisting of 164 original programming problems assessing language comprehension, algorithms, and simple mathematics
HumanEval-Average is a text benchmark evaluating models on reasoning tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.6.
Compare leaders on the best AI for reasoning leaderboards.
Current leaders
Codestral-22B from Mistral AI currently leads the HumanEval-Average leaderboard with a score of 0.615 across 1 evaluated AI models.
FAQ
Common questions about the HumanEval-Average benchmark and leaderboard.