The AI arena is free today

Open Superagent

SIFO

Progress Over Time

Interactive timeline showing model performance evolution on SIFO

State-of-the-art frontier
Open
Proprietary

SIFO Leaderboard

1 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B
Notice missing or incorrect data?

Sub-benchmarks

About this benchmark

What is SIFO?

SIFO (Simple Instruction Following) evaluates how well language models follow simple, explicit instructions. It tests fundamental instruction-following capabilities across various task types.

SIFO is a text benchmark evaluating models on instruction following, structured output, general, and agents tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–100 scale. The current average is 0.8, with the leader at 0.8.

Compare leaders on the best AI for instruction following, best AI for structured output, best AI for general and best AI for agents leaderboards.

Current leaders

Qwen3 VL 235B A22B Thinking from Alibaba Cloud / Qwen Team currently leads the SIFO leaderboard with a score of 0.773 across 1 evaluated AI models.

1Qwen3 VL 235B A22B ThinkingAlibaba Cloud / Qwen Team0.8%

FAQ

Common questions about the SIFO benchmark and leaderboard.

What is the SIFO benchmark?

SIFO (Simple Instruction Following) evaluates how well language models follow simple, explicit instructions. It tests fundamental instruction-following capabilities across various task types.

What is the SIFO leaderboard?

The SIFO leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Qwen3 VL 235B A22B Thinking by Alibaba Cloud / Qwen Team leads with a score of 0.773. The average score across all models is 0.773.

What is the highest SIFO score?

The highest SIFO score is 0.773, achieved by Qwen3 VL 235B A22B Thinking from Alibaba Cloud / Qwen Team.

How many models are evaluated on SIFO?

1 models have been evaluated on the SIFO benchmark, with 0 verified results and 1 self-reported results.

What categories does SIFO cover?

SIFO is categorized under instruction following, structured output, general, and agents. The benchmark evaluates text models.

Are there variants of SIFO?

Yes. SIFO has 1 related variant: SIFO-Multiturn.

What is the best open-source model on SIFO?

Qwen3 VL 235B A22B Thinking by Alibaba Cloud / Qwen Team is the top-ranked open-source model on SIFO, with a score of 0.773 (rank #1).

How recent are the SIFO leaderboard results?

The SIFO leaderboard was last updated in August 2026 and currently includes 1 evaluated models.
SIFO Leaderboard