SIFO-Multiturn
Progress Over Time
Interactive timeline showing model performance evolution on SIFO-Multiturn
SIFO-Multiturn Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 236B | — | — |
What is SIFO-Multiturn?
SIFO-Multiturn evaluates instruction following capabilities in multi-turn conversational settings, testing how well models maintain context and follow instructions across multiple exchanges.
SIFO-Multiturn is a text benchmark evaluating models on structured output, general, and agents tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–100 scale. The current average is 0.7, with the leader at 0.7.
Compare leaders on the best AI for structured output, best AI for general and best AI for agents leaderboards.
Current leaders
Qwen3 VL 235B A22B Thinking from Alibaba Cloud / Qwen Team currently leads the SIFO-Multiturn leaderboard with a score of 0.711 across 1 evaluated AI models.
FAQ
Common questions about the SIFO-Multiturn benchmark and leaderboard.