Home›Journal›This post

Sliding Window Attention: Map the Receptive Field

Map direct and transitive token reach through local and global attention layers, then publish exact edge and cache-position ledgers without inventing quality claims.

JP
JP Casabianca
AI Engineer and Product Designer · full-stack delivery · Bogotá

A 1,024-token attention window does not mean a deep model forgets everything 1,025 tokens away after one layer. It also does not prove the full context is equally usable. This sliding window attention tutorial maps the actual layer-by-layer dependency graph before anyone turns window size into a quality or memory claim.

Sliding window attention is not one number

A 1,024-token window describes the direct neighborhood offered by one local layer, not every path through a deep network. Consider a causal stack whose convention allows the current token plus four earlier positions. Token 12 can read tokens 8–12 in layer one. In layer two it can read the layer-one states at 8–12, and those states already contain paths back to token 4. After three local layers, token 0 can structurally influence token 12 even though those positions never share a direct edge.

That expansion is the receptive field: the transitive set of inputs with a directed path to the target representation. For a simple causal local window with four prior positions, the earliest reachable index after depth d is max(0, target − 4d). Change whether “window” includes the current token and the coefficient changes. Add a global layer and the graph can change in one step.

This distinction keeps sliding window attention claims honest. Reachability says information could travel; it does not say attention gives that information material weight, that training learned to preserve it, or that a model recalls it. RoPE scaling recall tests address empirical recall; this tutorial owns only the graph.

The first receipt therefore records direct range and transitive range separately. Quoting one advertised context length as though it were both is a category error.

Layered receptive coneA target token reaches progressively earlier positions through three causal local layers.DIRECT EDGES BECOME TRANSITIVE PATHSL0L1L2L3solid = direct · dashed = inherited · inclusive width = 5
A causal width-five fixture expands from token 11 to the whole 12-token prefix after three local layers.
Reachable source ranges
DepthReachableMeaning
011target state
17–11direct local keys
23–11transitive union
30–11whole prefix structurally reachable
Direct edge
One mask-permitted key-to-query relationship in one layer.
Transitive path
A source connected to the target through two or more layers.
  • Solid paths are direct.
  • Dashed paths are inherited through depth.

Freeze the sliding window attention contract

Before drawing an edge, write the sliding window attention mask contract. Declare whether attention is causal or bidirectional, whether the window count includes the query position, the sequence length, target token, ordered layer schedule, and every global token. These fields turn a diagram into a reproducible fixture.

This article uses an inclusive width: a causal width of five gives a query at index q the keys max(0, q − 4) through q. A bidirectional width of five uses radius two on each side. Even widths need an explicit left/right rule, so the downloadable lab chooses left=floor((w−1)/2) and right=w−1−left. It never silently imports another library’s convention.

Sequence boundaries truncate neighborhoods. Padding, packed examples, document masks, and prefix-LM rules can remove edges that a width formula would otherwise create. The lab models one unpadded sequence; a production receipt must add every active mask. Causality also fixes direction: an earlier key can influence a later query, not the reverse.

The ordered schedule matters because local and global operations do not commute. A local layer followed by a full layer produces a different intermediate graph than the reverse, even when the final target reaches the same positions. Context parallelism for long-context training should begin only after this layer contract is frozen, because sharding communication depends on the actual edges.

Build each layer as a directed graph

Represent every token state at layer l as a node. Add a directed edge from key position k at layer l−1 to query position q at layer l when the mask permits q to attend to k. Sliding window attention becomes concrete when every permitted relationship is an edge. Self edges matter because they carry the prior representation forward. Boundary truncation makes edge counts smaller near sequence ends.

For the 12-token causal fixture in the visual, full attention has 78 directed edges: query q owns q+1 keys, and the triangular sum is 12×13/2. An inclusive local width of five has 1+2+3+4+8×5=50 edges: the first four rows contain 1,2,3,4 edges and the remaining eight contain five. These are theoretical adjacency counts, not FLOPs, latency, or kernel measurements.

A global token needs two separate rules. “Everyone can attend to the global token” adds an inbound key to queries. “The global token attends everywhere” lets that token aggregate the sequence. Longformer-style task globals commonly use both directions, but a model implementation may not. A full global layer is different again: every query receives the complete permitted causal or bidirectional set.

The Longformer paper motivates local windows plus task-selected global attention. Use its pattern as a source, not as permission to assume every local-global architecture shares its mask.

Propagate attention reachability through depth

Start the target’s reachable set with itself. At each layer, replace every currently reachable query state with the union of its incoming key positions for that layer. This is boolean graph composition for sliding window attention: existence matters, while weights and activation magnitude do not. Store the set after every layer and the first layer where each source becomes reachable.

For target 11 in the 12-token fixture with inclusive causal width five, the ranges are 11 at depth zero, 7–11 after one layer, 3–11 after two, and 0–11 after three. A mutation that subtracts width instead of width−1 produces 6–11 after one layer and should fail the boundary test. A mutation that reports only the last layer’s neighbors produces 7–11 at every depth and confuses direct with transitive reach.

Boolean propagation scales safely under the lab’s cap of 512 tokens and 128 layers. A production compiler can use bitsets or sparse intervals, but optimization must preserve the reference result. Recording a canonical configuration beside the result makes replay more valuable than a screenshot.

Shortest influence depth is useful for review. It tells you how many transformations separate a source from the target. It still cannot establish that a usable feature survives those transformations. Structural attention reachability is a necessary-path statement, not a retrieval benchmark.

Attention-pattern triptychFull local and hub masks use distinct matrix shapes for the same twelve-token fixture.THE SAME 12 TOKENS · THREE EDGE SETSFULL · 78LOCAL · 50LOCAL + HUBtriangles encode causality; coral row and column encode a declared hub
Exact adjacency rules determine each matrix and edge count.
Pattern comparison
PatternEdgesTarget 11 direct keys
Full causal780–11
Local causal, width 5507–11
Local + hubfixture-dependentlocal keys plus hub

Model global layers and tokens explicitly

A global layer creates all mask-permitted edges for that layer. In the causal fixture, a global layer at depth two lets the target receive every earlier position immediately, regardless of what its local predecessor reached. In a bidirectional fixture it can connect the target to the full sequence. That discontinuity should appear as a bridge in the reach receipt, not be smoothed into an average window. This is a local-global attention contract, not merely a label.

Global tokens create hubs instead. If token 0 attends to every token and every token attends to token 0, a bidirectional local model can move information across distant regions in two or more steps. In a causal model, allowing token 0 to attend future positions would violate causality, so “global” must still obey the declared directional mask. The lab applies global-token expansion only where the selected mode permits it.

The Mistral 7B paper describes causal sliding-window attention and a rolling buffer for inference. The Gemma 3 technical report documents an explicit interleaving of local and global layers and discusses KV-cache trade-offs. Those examples show why the sliding window attention layer schedule belongs in the receipt; their measurements do not transfer to an unrelated model.

Use attention sinks for streaming LLMs for retained anchor behavior. A designated global token and a streaming sink are not interchangeable labels.

Separate reach, edge count, and cache policy

Three ledgers prevent one graph from claiming too much. The reach ledger lists sources that have a path to the target after each layer. The edge ledger counts permitted query-key pairs for the declared sequence. The cache ledger states which past positions an implementation says it retains during autoregressive decoding. None is a measured latency or quality result.

For a causal local window of width w at decode position q, a theoretical minimal local ledger may retain the last min(w,q+1) positions. The lab emits those exact indexes and their count for every scheduled layer at the declared target; a global layer lists the whole permitted prefix. Bidirectional mode reports cache retention as not applicable because this graph does not define autoregressive KV-cache semantics. Real kernels may allocate blocks, share storage, preallocate capacity, or keep extra positions, so a theoretical position count must not be relabeled bytes.

KV-cache optimization covers memory layouts and serving choices; the decision chart here uses only exact position and edge counts from one fixture. Similarly, LLM inference roofline analysis is the right place to connect operations and bandwidth to measured hardware behavior.

This boundary matters commercially. A smaller theoretical edge count can support an efficiency hypothesis, but only a benchmark on the served implementation can quantify speed or memory. Publish the formula, fixture, and observed metric separately so architecture reasoning remains useful when kernels change.

Reach edge and cache decision chartThree schedules compare target reach theoretical edges and retained-position declarations.STRUCTURE RECEIPT · NOT A BENCHMARKlocal, locallocal, globalfull, fullreach 9reach 12reach 12bar = reach after two layers; table carries edges and cache policy
Reachability, edge count, and cache policy remain separate columns.
Exact 12-token causal fixture
ScheduleReach after 2Layer edgesRetained positions
local, local3–11 (9)50 + 507–11 (5 positions)
local, global0–11 (12)50 + 780–11 (12 positions)
full, full0–11 (12)78 + 780–11 (12 positions)

Validate the architecture you actually serve

Begin with configuration and source inspection. Find the exact sliding window attention mask builder, inclusive-window convention, local/global schedule, global-token rule, cache eviction code, and padding behavior. Model cards are useful navigation, but executable configuration wins when prose and code diverge. Record the revision because a library update can change an off-by-one boundary. If documentation calls the mechanism windowed self-attention, verify that its edge rules match rather than treating the phrase as proof.

Then use adversarial fixtures. Test one token; a target at the left and right sequence edges; a window of one; the exact boundary at width−1; causal versus bidirectional mode; local-only expansion; a global layer before and after locals; a declared global token; duplicate globals; and the maximum accepted schedule. Mutation tests should prove the suite rejects an added boundary edge and a direct/transitive mix-up.

After structural parity, test the model. Place controlled evidence at distances that are unreachable, newly reachable, and reachable through many layers. Vary distractors and task wording. Reachable is not recalled, and an unreachable direct edge can still have a transitive path. Report model, weights, tokenizer, prompt template, hardware, decoding, repetitions, and uncertainty.

The local lab deliberately does not load a model or benchmark a kernel. Its job is smaller: make the architecture claim falsifiable before expensive evaluation. When the graph disagrees with the served implementation, fix the receipt first.

Publish a structural reachability receipt

A useful sliding window attention receipt names the sequence length, target, mode, window convention, ordered schedule, global tokens, direct neighborhood, per-layer reachable sets, first-reachable depth, theoretical edges, and the exact retained-position set and count for each applicable cache layer. Add a stable hash over canonical JSON so two reviewers can confirm they evaluated the same graph. Export the direct adjacency matrix as a standalone namespaced SVG as well as rendering it on the page; its rectangle count must equal the final layer’s edge ledger.

Also publish the limits: masks outside the fixture are not modeled; reach is possibility rather than weight, recall, quality, or causation; edge counts are not runtime measurements; and cache positions are a declared theoretical ledger. Revisit this article on 2027-01-28, or earlier when a cited configuration changes its local/global schedule, mask convention, or cache semantics.

The decision is not “local attention good or bad.” It is whether the architecture supplies a path required by the task, whether the served code matches that architecture, and whether model-specific evaluation shows the path is useful. That sequence separates structural reasoning from performance storytelling.

One page can now carry the whole claim: graph, formula, schedule, tests, model evaluation boundary, and refresh trigger. That is enough to replace a vague context-length promise with evidence another engineer can replay.

Runnable local artifact — Structural reach is possibility, not attention weight, trained recall, quality, causation, or measured runtime performance.

Plain text1 line
Declare tokens, target, mode, inclusive window, ordered local/global layers, and global tokens; calculate exact adjacency and propagate boolean reach.