The AI arena is free today

Open Superagent

IFEval

Paper

Progress Over Time

Interactive timeline showing model performance evolution on IFEval

State-of-the-art frontier
Open
Proprietary

IFEval Leaderboard

68 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.30 / $2.40
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
5
OpenAI
OpenAI
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B
7
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B
970B
9
Amazon
Amazon
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
12
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
1327B
149B
154B
16
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B
161.0T
16
Moonshot AI
Moonshot AI
1.0T
19
Amazon
Amazon
20560B
21
LG AI Research
LG AI Research
33B
22253B
23
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
80B
2312B
25
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
235B
26405B
27
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B
27
OpenAI
OpenAI
29
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B
29
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
235B
29
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B
32
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
80B
3370B
34
OpenAI
OpenAI
1.0M$2.00 / $8.00
35
Moonshot AI
Moonshot AI
35
37
DeepSeek
DeepSeek
671B
38
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
31B
3914B
40
Sarvam AI
Sarvam AI
105B
41
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B
421.0M$0.40 / $1.60
42
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
73B
44
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B
45
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
4614B
47
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
4824B
49
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B262K$0.10 / $1.00
50
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B262K$0.10 / $0.60
150 of 68
1/2
Notice missing or incorrect data?
About this benchmark

What is IFEval?

Instruction-Following Evaluation (IFEval) benchmark for large language models, focusing on verifiable instructions with 25 types of instructions and around 500 prompts containing one or more verifiable constraints

IFEval is a text benchmark evaluating models on chat, structured output, instruction following, and general tasks. LLM Stats tracks 68 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.9.

Compare leaders on the best AI for chat, best AI for structured output, best AI for instruction following and best AI for general leaderboards.

Current leaders

Qwen3.5-27B from Alibaba Cloud / Qwen Team currently leads the IFEval leaderboard with a score of 0.950 across 68 evaluated AI models.

1Qwen3.5-27BAlibaba Cloud / Qwen Team95.0%
2Qwen3.7-PlusAlibaba Cloud / Qwen Team94.6%
3Qwen3.7 MaxAlibaba Cloud / Qwen Team94.3%

Source paper

Title
Instruction-Following Evaluation for Large Language Models
Authors
Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, and 4 others
Published
Abstract

One core capability of Large Language Models (LLMs) is to follow natural language instructions. However, the evaluation of such abilities is not standardized: Human evaluations are expensive, slow, and not objectively reproducible, while LLM-based auto-evaluation is potentially biased or limited by the ability of the evaluator LLM. To overcome these issues, we introduce Instruction-Following Eval (IFEval) for large language models. IFEval is a straightforward and easy-to-reproduce evaluation benchmark. It focuses on a set of "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times". We identified 25 types of those verifiable instructions and constructed around 500 prompts, with each prompt containing one or more verifiable instructions. We show evaluation results of two widely available LLMs on the market. Our code and data can be found at https://github.com/google-research/google-research/tree/master/instruction_following_eval

FAQ

Common questions about the IFEval benchmark and leaderboard.

What is the IFEval benchmark?

Instruction-Following Evaluation (IFEval) benchmark for large language models, focusing on verifiable instructions with 25 types of instructions and around 500 prompts containing one or more verifiable constraints

What is the IFEval leaderboard?

The IFEval leaderboard ranks 68 AI models based on their performance on this benchmark. Currently, Qwen3.5-27B by Alibaba Cloud / Qwen Team leads with a score of 0.950. The average score across all models is 0.845.

What is the highest IFEval score?

The highest IFEval score is 0.950, achieved by Qwen3.5-27B from Alibaba Cloud / Qwen Team.

How many models are evaluated on IFEval?

68 models have been evaluated on the IFEval benchmark, with 0 verified results and 68 self-reported results.

Where can I find the IFEval paper?

The IFEval paper is available at https://arxiv.org/abs/2311.07911. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does IFEval cover?

IFEval is categorized under chat, structured output, instruction following, and general. The benchmark evaluates text models.

What is the best open-source model on IFEval?

Qwen3.5-27B by Alibaba Cloud / Qwen Team is the top-ranked open-source model on IFEval, with a score of 0.950 (rank #1).

Which model offers the best value on IFEval?

Among models scoring within 10% of the leader, Qwen3.5-27B from Alibaba Cloud / Qwen Team is the cheapest, at $0.30 per million input tokens with a score of 0.950.

How recent are the IFEval leaderboard results?

The IFEval leaderboard was last updated in September 2026 and currently includes 68 evaluated models.