ModelRefs / HellaSwag Methodology — Methodology
HellaSwag Methodology — Methodology
HellaSwag tests commonsense sentence completion with adversarially filtered distractors.
Overview
What it measures: Commonsense plausibility — picking the most natural continuation of a short context.
How it works
- ~10K validation items.
- Each item: a context and 4 candidate completions.
- Distractors filtered to be hard for prior models.
Strengths
Cheap and well-established
Limitations
- Saturated
- Limited reasoning depth
Best use cases
Sanity-check eval for small models
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to HellaSwag Methodology — Methodology.
Frequently asked questions
What does HellaSwag measure?
Commonsense plausibility — picking the most natural continuation of a short context.
What are its main limitations?
Saturated Limited reasoning depth
When should I use this benchmark?
Sanity-check eval for small models