ModelRefs / Sliding Window Attention — AI Glossary
Sliding Window Attention — AI Glossary
An attention pattern restricting each token to attend only to a local window of neighbors, enabling efficient long-context inference.
Overview
Sliding window attention (Beltagy et al. 2020; used in Mistral) limits the attention field to W tokens per position, reducing quadratic complexity to O(n·W). Global tokens or interleaved full-attention layers handle long-range dependencies. Enables models trained on shorter contexts to infer over longer sequences efficiently.
Reference details
| Topic | architecture |
|---|---|
| Also known as | local attention, SWA |
| Last reviewed | 2026-06-24 |
Related terms
Primary source
Continue your research
Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Sliding Window Attention — AI Glossary.
Frequently asked questions
What is Sliding Window Attention?
An attention pattern restricting each token to attend only to a local window of neighbors, enabling efficient long-context inference.
Is Sliding Window Attention the same as local attention?
Yes — local attention, SWA are common aliases for Sliding Window Attention.
What concepts are related to Sliding Window Attention?
Closely related concepts include attention, long context, grouped query attention.