ModelRefs / WinoGrande Leaderboard — AI Model Scores
WinoGrande Leaderboard — AI Model Scores
Large-scale Winograd-schema commonsense pronoun resolution (WinoGrande XL). Current leaders, methodology, and citation sources for WinoGrande.
Overview
Large-scale Winograd-schema commonsense pronoun resolution (WinoGrande XL).
How it is measured: Binary pronoun selection accuracy on the xl split (1267 questions); debiased with AFLITE.
How this benchmark is scored
| Category | open-source |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://winogrande.allenai.org/
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to WinoGrande Leaderboard — AI Model Scores.