ModelRefs / Vocabulary Size — AI Glossary

Vocabulary Size — AI Glossary

The number of distinct tokens in a model's tokenizer; larger vocabularies improve multilingual coverage at the cost of embedding table memory.

Overview

GPT-2 used 50,257 tokens; GPT-4 cl100k_base has 100,277; Llama 3 uses 128,256. Larger vocabularies reduce average tokens per word for non-English text (fewer characters per token), improving multilingual efficiency. The embedding matrix grows with vocabulary size, adding memory overhead.

Reference details

Topicarchitecture
Also known asvocab size, tokenizer vocabulary
Last reviewed2026-06-24

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Vocabulary Size — AI Glossary.

Frequently asked questions

What is Vocabulary Size?

The number of distinct tokens in a model's tokenizer; larger vocabularies improve multilingual coverage at the cost of embedding table memory.

Is Vocabulary Size the same as vocab size?

Yes — vocab size, tokenizer vocabulary are common aliases for Vocabulary Size.

What concepts are related to Vocabulary Size?

Closely related concepts include byte pair encoding, tiktoken, embedding dimension.