ModelRefs / Sliding Window Attention — AI Glossary

Sliding Window Attention — AI Glossary

An attention pattern restricting each token to attend only to a local window of neighbors, enabling efficient long-context inference.

Overview

Sliding window attention (Beltagy et al. 2020; used in Mistral) limits the attention field to W tokens per position, reducing quadratic complexity to O(n·W). Global tokens or interleaved full-attention layers handle long-range dependencies. Enables models trained on shorter contexts to infer over longer sequences efficiently.

Reference details

Topicarchitecture
Also known aslocal attention, SWA
Last reviewed2026-06-24

Primary source

Continue your research

Use these connected ModelRefs sections to compare alternatives, inspect implementation paths, and review the evidence and governance boundaries relevant to Sliding Window Attention — AI Glossary.

Frequently asked questions

What is Sliding Window Attention?

An attention pattern restricting each token to attend only to a local window of neighbors, enabling efficient long-context inference.

Is Sliding Window Attention the same as local attention?

Yes — local attention, SWA are common aliases for Sliding Window Attention.

What concepts are related to Sliding Window Attention?

Closely related concepts include attention, long context, grouped query attention.