ModelRefs / GSM8K — AI Glossary
GSM8K — AI Glossary
A benchmark of 8,500 grade-school math word problems requiring multi-step arithmetic reasoning. GPT-4 scores ~97%; smaller 7B fine-tuned models reach 70–80%.
Overview
GSM8K (Cobbe et al. 2021) tests multi-step elementary math: addition, subtraction, multiplication, fractions. GPT-3 scored 35%; chain-of-thought prompting (Kojima et al.) increased scores to 60%. GPT-4 scores ~97%; smaller 7B fine-tuned models reach 70–80%. Near-saturated for frontier models; graduated to MATH and competition math.
Reference details
| Topic | evaluation |
|---|---|
| Last reviewed | 2026-06-24 |
Related terms
Primary source
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GSM8K — AI Glossary.
Frequently asked questions
What is GSM8K?
A benchmark of 8,500 grade-school math word problems requiring multi-step arithmetic reasoning.
What concepts are related to GSM8K?
Closely related concepts include math benchmark, chain of thought, reasoning.