ModelRefs / BIG-Bench Hard Leaderboard — AI Model Scores
BIG-Bench Hard Leaderboard — AI Model Scores
23 challenging tasks from BIG-Bench where prior LMs underperformed humans. Current leaders, methodology, and citation sources for BIG-Bench Hard.
Overview
23 challenging tasks from BIG-Bench where prior LMs underperformed humans.
How it is measured: 3-shot CoT; reported as macro-average accuracy.
How this benchmark is scored
| Category | reasoning |
|---|---|
| Maximum score | 100 % accuracy |
| Direction | Higher is better |
Primary source: https://arxiv.org/abs/2210.09261
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to BIG-Bench Hard Leaderboard — AI Model Scores.