ModelRefs / MuSR Leaderboard — AI Model Scores

MuSR Leaderboard — AI Model Scores

Multistep soft reasoning over narrative scenarios. Current leaders, methodology, and citation sources for MuSR.

Overview

Multistep soft reasoning over narrative scenarios.

How it is measured: Zero-shot accuracy on murder mystery / team allocation / object placements.

How this benchmark is scored

Categoryreasoning
Maximum score100 % accuracy
DirectionHigher is better

Primary source: https://arxiv.org/abs/2310.16049

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to MuSR Leaderboard — AI Model Scores.