SIFO
Progress Over Time
Interactive timeline showing model performance evolution on SIFO
SIFO Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 236B | — | — |
Sub-benchmarks
What is SIFO?
SIFO (Simple Instruction Following) evaluates how well language models follow simple, explicit instructions. It tests fundamental instruction-following capabilities across various task types.
SIFO is a text benchmark evaluating models on instruction following, structured output, general, and agents tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–100 scale. The current average is 0.8, with the leader at 0.8.
Compare leaders on the best AI for instruction following, best AI for structured output, best AI for general and best AI for agents leaderboards.
Current leaders
Qwen3 VL 235B A22B Thinking from Alibaba Cloud / Qwen Team currently leads the SIFO leaderboard with a score of 0.773 across 1 evaluated AI models.
FAQ
Common questions about the SIFO benchmark and leaderboard.