ModelRefs / Vocabulary Size — AI Glossary
Vocabulary Size — AI Glossary
The number of distinct tokens in a model's tokenizer; larger vocabularies improve multilingual coverage at the cost of embedding table memory.
Overview
GPT-2 used 50,257 tokens; GPT-4 cl100k_base has 100,277; Llama 3 uses 128,256. Larger vocabularies reduce average tokens per word for non-English text (fewer characters per token), improving multilingual efficiency. The embedding matrix grows with vocabulary size, adding memory overhead.
Reference details
| Topic | architecture |
|---|---|
| Also known as | vocab size, tokenizer vocabulary |
| Last reviewed | 2026-06-24 |
Related terms
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Vocabulary Size — AI Glossary.
Frequently asked questions
What is Vocabulary Size?
The number of distinct tokens in a model's tokenizer; larger vocabularies improve multilingual coverage at the cost of embedding table memory.
Is Vocabulary Size the same as vocab size?
Yes — vocab size, tokenizer vocabulary are common aliases for Vocabulary Size.
What concepts are related to Vocabulary Size?
Closely related concepts include byte pair encoding, tiktoken, embedding dimension.