Multilingual MGSM (CoT)
Progress Over Time
Interactive timeline showing model performance evolution on Multilingual MGSM (CoT)
Multilingual MGSM (CoT) Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | 405B | — | — | |||
| 2 | 70B | 131K | $0.40 / $0.40 | |||
| 3 | 8B | 131K | $0.02 / $0.04 |
What is Multilingual MGSM (CoT)?
Multilingual Grade School Math (MGSM) benchmark evaluates language models' chain-of-thought reasoning abilities across ten typologically diverse languages. Contains 250 grade-school math problems manually translated from GSM8K dataset into languages including Bengali and Swahili.
Multilingual MGSM (CoT) is a text benchmark evaluating models on math and reasoning tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.9.
Compare leaders on the best AI for math and best AI for reasoning leaderboards.
Current leaders
Llama 3.1 405B Instruct from Meta currently leads the Multilingual MGSM (CoT) leaderboard with a score of 0.916 across 3 evaluated AI models.
FAQ
Common questions about the Multilingual MGSM (CoT) benchmark and leaderboard.