ModelRefs / GSM8K — AI Glossary

GSM8K — AI Glossary

A benchmark of 8,500 grade-school math word problems requiring multi-step arithmetic reasoning. GPT-4 scores ~97%; smaller 7B fine-tuned models reach 70–80%.

Overview

GSM8K (Cobbe et al. 2021) tests multi-step elementary math: addition, subtraction, multiplication, fractions. GPT-3 scored 35%; chain-of-thought prompting (Kojima et al.) increased scores to 60%. GPT-4 scores ~97%; smaller 7B fine-tuned models reach 70–80%. Near-saturated for frontier models; graduated to MATH and competition math.

Reference details

Topicevaluation
Last reviewed2026-06-24

Primary source

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to GSM8K — AI Glossary.

Frequently asked questions

What is GSM8K?

A benchmark of 8,500 grade-school math word problems requiring multi-step arithmetic reasoning.

What concepts are related to GSM8K?

Closely related concepts include math benchmark, chain of thought, reasoning.