The AI arena is free today

Open Superagent

IFBench

Progress Over Time

Interactive timeline showing model performance evolution on IFBench

State-of-the-art frontier
Open
Proprietary

IFBench Leaderboard

42 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
2
Thinking Machines Lab
Thinking Machines Lab
276B524K$0.30 / $1.20
3550B262K$0.50 / $2.20
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B1.0M$0.15 / $0.47
6
Nous Research
Nous Research
70B131K$0.70 / $0.70
7
8
Thinking Machines Lab
Thinking Machines Lab
975B524K$0.95 / $4.05
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.40 / $3.00
108B131K$0.06 / $0.25
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
1330B131K$0.16 / $0.65
1430B131K$0.30 / $1.20
15
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.26 / $2.60
15
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B262K$0.45 / $3.00
17
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B262K$0.29 / $2.40
18
19
InclusionAI
InclusionAI
124B131K$0.06 / $0.18
203B131K$0.03 / $0.12
21
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
22218B
23120B262K$0.09 / $0.40
2430B262K$0.08 / $0.20
25
Inception
Inception
128K$0.25 / $0.75
26
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0T256K$1.20 / $6.00
271.0M$0.30 / $2.50
28
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.14 / $1.00
29230B1.0M$0.30 / $1.20
30117B
311.0T
31128B256K$1.50 / $7.50
33
34
LG AI Research
LG AI Research
236B
35
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B262K$0.10 / $0.15
36
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B
37
Liquid AI
Liquid AI
3B
38
Mistral AI
Mistral AI
119B256K$0.15 / $0.60
39
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2B
401.0M$0.33 / $2.75
41
Liquid AI
Liquid AI
3B
42
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
800M
Notice missing or incorrect data?
About this benchmark

What is IFBench?

Instruction Following Benchmark evaluating model's ability to follow complex instructions

IFBench is a text benchmark evaluating models on instruction following and general tasks. LLM Stats tracks 42 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.8.

Compare leaders on the best AI for instruction following and best AI for general leaderboards.

Current leaders

Qwen3.8 Max from Alibaba Cloud / Qwen Team currently leads the IFBench leaderboard with a score of 0.828 across 42 evaluated AI models.

1Qwen3.8 MaxAlibaba Cloud / Qwen Team82.8%
2Inkling-SmallThinking Machines Lab82.2%

FAQ

Common questions about the IFBench benchmark and leaderboard.

What is the IFBench benchmark?

Instruction Following Benchmark evaluating model's ability to follow complex instructions

What is the IFBench leaderboard?

The IFBench leaderboard ranks 42 AI models based on their performance on this benchmark. Currently, Qwen3.8 Max by Alibaba Cloud / Qwen Team leads with a score of 0.828. The average score across all models is 0.695.

What is the highest IFBench score?

The highest IFBench score is 0.828, achieved by Qwen3.8 Max from Alibaba Cloud / Qwen Team.

How many models are evaluated on IFBench?

42 models have been evaluated on the IFBench benchmark, with 0 verified results and 42 self-reported results.

What categories does IFBench cover?

IFBench is categorized under instruction following and general. The benchmark evaluates text models.

What is the best open-source model on IFBench?

Inkling-Small by Thinking Machines Lab is the top-ranked open-source model on IFBench, with a score of 0.822 (rank #2).

Which model offers the best value on IFBench?

Among models scoring within 10% of the leader, IBM Granite 4.2 8B from IBM is the cheapest, at $0.06 per million input tokens with a score of 0.793.

How recent are the IFBench leaderboard results?

The IFBench leaderboard was last updated in September 2026 and currently includes 42 evaluated models.