The AI arena is free today

Open Superagent

MuSR

Paper

Progress Over Time

Interactive timeline showing model performance evolution on MuSR

State-of-the-art frontier
Open
Proprietary

MuSR Leaderboard

2 models
ContextCostLicense
1
Moonshot AI
Moonshot AI
1.0T
2
Nous Research
Nous Research
70B
Notice missing or incorrect data?
About this benchmark

What is MuSR?

MuSR (Multistep Soft Reasoning) is a benchmark for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Created through a neurosymbolic synthetic-to-natural generation algorithm, it generates complex reasoning scenarios like murder mysteries roughly 1000 words in length that challenge current LLMs including GPT-4. The benchmark tests chain-of-thought reasoning capabilities across domains involving commonsense reasoning about physical and social situations.

MuSR is a text benchmark evaluating models on reasoning tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.8.

Compare leaders on the best AI for reasoning leaderboards.

Current leaders

Kimi K2 Instruct from Moonshot AI currently leads the MuSR leaderboard with a score of 0.764 across 2 evaluated AI models.

1Kimi K2 InstructMoonshot AI76.4%
2Hermes 3 70BNous Research50.7%

Source paper

Title
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
Authors
Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and 1 others
Published
Abstract

While large language models (LLMs) equipped with techniques like chain-of-thought prompting have demonstrated impressive capabilities, they still fall short in their ability to reason robustly in complex settings. However, evaluating LLM reasoning is challenging because system capabilities continue to grow while benchmark datasets for tasks like logical deduction have remained static. We introduce MuSR, a dataset for evaluating language models on multistep soft reasoning tasks specified in a natural language narrative. This dataset has two crucial features. First, it is created through a novel neurosymbolic synthetic-to-natural generation algorithm, enabling the construction of complex reasoning instances that challenge GPT-4 (e.g., murder mysteries roughly 1000 words in length) and which can be scaled further as more capable LLMs are released. Second, our dataset instances are free text narratives corresponding to real-world domains of reasoning; this makes it simultaneously much more challenging than other synthetically-crafted benchmarks while remaining realistic and tractable for human annotators to solve with high accuracy. We evaluate a range of LLMs and prompting techniques on this dataset and characterize the gaps that remain for techniques like chain-of-thought to perform robust reasoning.

FAQ

Common questions about the MuSR benchmark and leaderboard.

What is the MuSR benchmark?

MuSR (Multistep Soft Reasoning) is a benchmark for evaluating language models on multistep soft reasoning tasks specified in natural language narratives. Created through a neurosymbolic synthetic-to-natural generation algorithm, it generates complex reasoning scenarios like murder mysteries roughly 1000 words in length that challenge current LLMs including GPT-4. The benchmark tests chain-of-thought reasoning capabilities across domains involving commonsense reasoning about physical and social situations.

What is the MuSR leaderboard?

The MuSR leaderboard ranks 2 AI models based on their performance on this benchmark. Currently, Kimi K2 Instruct by Moonshot AI leads with a score of 0.764. The average score across all models is 0.635.

What is the highest MuSR score?

The highest MuSR score is 0.764, achieved by Kimi K2 Instruct from Moonshot AI.

How many models are evaluated on MuSR?

2 models have been evaluated on the MuSR benchmark, with 0 verified results and 2 self-reported results.

Where can I find the MuSR paper?

The MuSR paper is available at https://arxiv.org/abs/2310.16049. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does MuSR cover?

MuSR is categorized under reasoning. The benchmark evaluates text models.

What is the best open-source model on MuSR?

Kimi K2 Instruct by Moonshot AI is the top-ranked open-source model on MuSR, with a score of 0.764 (rank #1).

How recent are the MuSR leaderboard results?

The MuSR leaderboard was last updated in September 2026 and currently includes 2 evaluated models.