IFBench
Progress Over Time
Interactive timeline showing model performance evolution on IFBench
State-of-the-art frontier
Open
Proprietary
IFBench Leaderboard
42 models
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 2.4T | 1.0M | $1.65 / $4.95 | ||
| 2 | Thinking Machines Lab | 276B | 524K | $0.30 / $1.20 | ||
| 3 | 550B | 262K | $0.50 / $2.20 | |||
| 4 | Alibaba Cloud / Qwen Team | 125B | — | — | ||
| 4 | Alibaba Cloud / Qwen Team | 125B | 1.0M | $0.15 / $0.47 | ||
| 6 | Nous Research | 70B | 131K | $0.70 / $0.70 | ||
| 7 | Amazon | — | — | — | ||
| 8 | Thinking Machines Lab | 975B | 524K | $0.95 / $4.05 | ||
| 9 | Alibaba Cloud / Qwen Team | 28B | 262K | $0.40 / $3.00 | ||
| 10 | 8B | 131K | $0.06 / $0.25 | |||
| 11 | Alibaba Cloud / Qwen Team | — | — | — | ||
| 11 | Alibaba Cloud / Qwen Team | — | 1.0M | $1.25 / $3.75 | ||
| 13 | 30B | 131K | $0.16 / $0.65 | |||
| 14 | Meta | 30B | 131K | $0.30 / $1.20 | ||
| 15 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.26 / $2.60 | ||
| 15 | Alibaba Cloud / Qwen Team | 397B | 262K | $0.45 / $3.00 | ||
| 17 | Alibaba Cloud / Qwen Team | 122B | 262K | $0.29 / $2.40 | ||
| 18 | Microsoft | — | — | — | ||
| 19 | InclusionAI | 124B | 131K | $0.06 / $0.18 | ||
| 20 | 3B | 131K | $0.03 / $0.12 | |||
| 21 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 22 | Cohere | 218B | — | — | ||
| 23 | 120B | 262K | $0.09 / $0.40 | |||
| 24 | 30B | 262K | $0.08 / $0.20 | |||
| 25 | Inception | — | 128K | $0.25 / $0.75 | ||
| 26 | Alibaba Cloud / Qwen Team | 1.0T | 256K | $1.20 / $6.00 | ||
| 27 | Amazon | — | 1.0M | $0.30 / $2.50 | ||
| 28 | Alibaba Cloud / Qwen Team | 35B | 262K | $0.14 / $1.00 | ||
| 29 | MiniMax | 230B | 1.0M | $0.30 / $1.20 | ||
| 30 | OpenAI | 117B | — | — | ||
| 31 | Microsoft | 1.0T | — | — | ||
| 31 | Mistral AI | 128B | 256K | $1.50 / $7.50 | ||
| 33 | Amazon | — | — | — | ||
| 34 | LG AI Research | 236B | — | — | ||
| 35 | Alibaba Cloud / Qwen Team | 9B | 262K | $0.10 / $0.15 | ||
| 36 | Alibaba Cloud / Qwen Team | 4B | — | — | ||
| 37 | Liquid AI | 3B | — | — | ||
| 38 | Mistral AI | 119B | 256K | $0.15 / $0.60 | ||
| 39 | Alibaba Cloud / Qwen Team | 2B | — | — | ||
| 40 | Amazon | — | 1.0M | $0.33 / $2.75 | ||
| 41 | Liquid AI | 3B | — | — | ||
| 42 | Alibaba Cloud / Qwen Team | 800M | — | — |
Notice missing or incorrect data?
What is IFBench?
Instruction Following Benchmark evaluating model's ability to follow complex instructions
IFBench is a text benchmark evaluating models on instruction following and general tasks. LLM Stats tracks 42 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.8.
Compare leaders on the best AI for instruction following and best AI for general leaderboards.
Current leaders
Qwen3.8 Max from Alibaba Cloud / Qwen Team currently leads the IFBench leaderboard with a score of 0.828 across 42 evaluated AI models.
FAQ
Common questions about the IFBench benchmark and leaderboard.