IFEval
Progress Over Time
Interactive timeline showing model performance evolution on IFEval
IFEval Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.30 / $2.40 | ||
| 2 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.32 / $1.28 | ||
| 3 | Alibaba Cloud / Qwen Team | — | 1.0M | $1.25 / $3.75 | ||
| 3 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 5 | OpenAI | — | — | — | ||
| 6 | Alibaba Cloud / Qwen Team | 122B | — | — | ||
| 7 | Anthropic | — | — | — | ||
| 8 | Alibaba Cloud / Qwen Team | 397B | — | — | ||
| 9 | Amazon | — | — | — | ||
| 9 | 70B | — | — | |||
| 11 | Alibaba Cloud / Qwen Team | 35B | — | — | ||
| 12 | Alibaba Cloud / Qwen Team | 9B | — | — | ||
| 13 | Google | 27B | — | — | ||
| 14 | NVIDIA | 9B | — | — | ||
| 15 | Google | 4B | — | — | ||
| 16 | Moonshot AI | 1.0T | — | — | ||
| 16 | Moonshot AI | 1.0T | — | — | ||
| 16 | Alibaba Cloud / Qwen Team | 4B | — | — | ||
| 19 | Amazon | — | — | — | ||
| 20 | Meituan | 560B | — | — | ||
| 21 | 253B | — | — | |||
| 22 | Alibaba Cloud / Qwen Team | 80B | — | — | ||
| 22 | Google | 12B | — | — | ||
| 24 | Alibaba Cloud / Qwen Team | 235B | — | — | ||
| 25 | 405B | — | — | |||
| 26 | OpenAI | — | — | — | ||
| 26 | Alibaba Cloud / Qwen Team | 236B | — | — | ||
| 28 | Alibaba Cloud / Qwen Team | 33B | — | — | ||
| 28 | Alibaba Cloud / Qwen Team | 236B | — | — | ||
| 28 | Alibaba Cloud / Qwen Team | 235B | — | — | ||
| 31 | Alibaba Cloud / Qwen Team | 80B | — | — | ||
| 32 | 70B | — | — | |||
| 33 | OpenAI | — | 1.0M | $2.00 / $8.00 | ||
| 34 | Amazon | — | — | — | ||
| 34 | Moonshot AI | — | — | — | ||
| 36 | DeepSeek | 671B | — | — | ||
| 37 | Alibaba Cloud / Qwen Team | 31B | — | — | ||
| 38 | Microsoft | 14B | — | — | ||
| 39 | Sarvam AI | 105B | — | — | ||
| 40 | Alibaba Cloud / Qwen Team | 33B | — | — | ||
| 41 | OpenAI | — | 1.0M | $0.40 / $1.60 | ||
| 41 | Alibaba Cloud / Qwen Team | 73B | — | — | ||
| 43 | Alibaba Cloud / Qwen Team | 33B | — | — | ||
| 44 | Alibaba Cloud / Qwen Team | 9B | — | — | ||
| 45 | Microsoft | 14B | — | — | ||
| 46 | Alibaba Cloud / Qwen Team | 9B | — | — | ||
| 47 | Mistral AI | 24B | — | — | ||
| 48 | Alibaba Cloud / Qwen Team | 4B | 262K | $0.10 / $1.00 | ||
| 49 | Alibaba Cloud / Qwen Team | 4B | 262K | $0.10 / $0.60 | ||
| 50 | Alibaba Cloud / Qwen Team | 31B | — | — |
What is IFEval?
Instruction-Following Evaluation (IFEval) benchmark for large language models, focusing on verifiable instructions with 25 types of instructions and around 500 prompts containing one or more verifiable constraints
IFEval is a text benchmark evaluating models on structured output, instruction following, and general tasks. LLM Stats tracks 65 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.9.
Compare leaders on the best AI for structured output, best AI for instruction following and best AI for general leaderboards.
Current leaders
Qwen3.5-27B from Alibaba Cloud / Qwen Team currently leads the IFEval leaderboard with a score of 0.950 across 65 evaluated AI models.
Source paper
- Title
- Instruction-Following Evaluation for Large Language Models
- Authors
- Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, and 4 others
- Published
- arXiv
- 2311.07911
Abstract
One core capability of Large Language Models (LLMs) is to follow natural language instructions. However, the evaluation of such abilities is not standardized: Human evaluations are expensive, slow, and not objectively reproducible, while LLM-based auto-evaluation is potentially biased or limited by the ability of the evaluator LLM. To overcome these issues, we introduce Instruction-Following Eval (IFEval) for large language models. IFEval is a straightforward and easy-to-reproduce evaluation benchmark. It focuses on a set of "verifiable instructions" such as "write in more than 400 words" and "mention the keyword of AI at least 3 times". We identified 25 types of those verifiable instructions and constructed around 500 prompts, with each prompt containing one or more verifiable instructions. We show evaluation results of two widely available LLMs on the market. Our code and data can be found at https://github.com/google-research/google-research/tree/master/instruction_following_eval
FAQ
Common questions about the IFEval benchmark and leaderboard.