IFEval

Paper

Progress Over Time

Interactive timeline showing model performance evolution on IFEval

State-of-the-art frontier
Open
Proprietary

IFEval Leaderboard

65 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.30 / $2.40
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.32 / $1.28
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
5
OpenAI
OpenAI
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B
7
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B
9
Amazon
Amazon
970B
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
12
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
1327B
149B
154B
16
Moonshot AI
Moonshot AI
1.0T
161.0T
16
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B
19
Amazon
Amazon
20560B
21253B
22
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
80B
2212B
24
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
235B
25405B
26
OpenAI
OpenAI
26
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B
28
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B
28
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B
28
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
235B
31
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
80B
3270B
33
OpenAI
OpenAI
1.0M$2.00 / $8.00
34
34
Moonshot AI
Moonshot AI
36
DeepSeek
DeepSeek
671B
37
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
31B
3814B
39
Sarvam AI
Sarvam AI
105B
40
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B
411.0M$0.40 / $1.60
41
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
73B
43
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B
44
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
4514B
46
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
4724B
48
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B262K$0.10 / $1.00
49
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B262K$0.10 / $0.60
50
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
31B
150 of 65
1/2
Notice missing or incorrect data?
About this benchmark

What is IFEval?

Instruction-Following Evaluation (IFEval) benchmark for large language models, focusing on verifiable instructions with 25 types of instructions and around 500 prompts containing one or more verifiable constraints

IFEval is a text benchmark evaluating models on structured output, instruction following, and general tasks. LLM Stats tracks 65 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.9.

Compare leaders on the best AI for structured output, best AI for instruction following and best AI for general leaderboards.

Current leaders

Qwen3.5-27B from Alibaba Cloud / Qwen Team currently leads the IFEval leaderboard with a score of 0.950 across 65 evaluated AI models.

1Qwen3.5-27BAlibaba Cloud / Qwen Team95.0%
2Qwen3.7-PlusAlibaba Cloud / Qwen Team94.6%
3Qwen3.7 MaxAlibaba Cloud / Qwen Team94.3%

Source paper

Title
Instruction-Following Evaluation for Large Language Models
Authors
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, and 4 others
Published
Abstract

One core capability of Large Language Models (LLMs) is to follow natural language instructions. However, the evaluation of such abilities is not standardized: Human evaluations are expensive, slow, and not objectively reproducible, while LLM-based auto-evaluation is potentially biased or limited by the ability of the evaluator LLM. To overcome these issues, we introduce Instruction-Following Eval (IFEval) for large language models. IFEval is a straightforward and easy-to-reproduce evaluation benchmark. It focuses on a set of "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times". We identified 25 types of those verifiable instructions and constructed around 500 prompts, with each prompt containing one or more verifiable instructions. We show evaluation results of two widely available LLMs on the market. Our code and data can be found at https://github.com/google-research/google-research/tree/master/instruction_following_eval

FAQ

Common questions about the IFEval benchmark and leaderboard.

What is the IFEval benchmark?

Instruction-Following Evaluation (IFEval) benchmark for large language models, focusing on verifiable instructions with 25 types of instructions and around 500 prompts containing one or more verifiable constraints

What is the IFEval leaderboard?

The IFEval leaderboard ranks 65 AI models based on their performance on this benchmark. Currently, Qwen3.5-27B by Alibaba Cloud / Qwen Team leads with a score of 0.950. The average score across all models is 0.846.

What is the highest IFEval score?

The highest IFEval score is 0.950, achieved by Qwen3.5-27B from Alibaba Cloud / Qwen Team.

How many models are evaluated on IFEval?

65 models have been evaluated on the IFEval benchmark, with 0 verified results and 65 self-reported results.

Where can I find the IFEval paper?

The IFEval paper is available at https://arxiv.org/abs/2311.07911. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does IFEval cover?

IFEval is categorized under structured output, instruction following, and general. The benchmark evaluates text models.

What is the best open-source model on IFEval?

Qwen3.5-27B by Alibaba Cloud / Qwen Team is the top-ranked open-source model on IFEval, with a score of 0.950 (rank #1).

Which model offers the best value on IFEval?

Among models scoring within 10% of the leader, Qwen3.5-27B from Alibaba Cloud / Qwen Team is the cheapest, at $0.30 per million input tokens with a score of 0.950.

How recent are the IFEval leaderboard results?

The IFEval leaderboard was last updated in July 2026 and currently includes 65 evaluated models.