LLM Request Preemption: Recompute or Swap?
Model one KV-cache pressure event, preserve request progress, compare recompute with two-way swap traffic, and expose starvation before rollout.
LLM request preemption is the moment a serving engine trades one request's progress for enough KV-cache space to keep the batch moving. This guide turns that invisible recovery cost into a replayable event ledger so an engineer can compare recomputation, swap traffic, queue delay, and starvation without pretending a browser model is a GPU benchmark.
LLM request preemption starts after admission
LLM request preemption happens after a request has already earned admission, consumed KV-cache blocks, and begun producing output. Admission control asks whether new work may enter. Preemption asks which running sequence must surrender memory when the active set no longer fits. Mixing those decisions makes a warning count difficult to interpret because a conservative admission limit can hide a poor recovery policy, while an aggressive limit can manufacture recovery work.
The pressure comes from uncertain growth. Prompt length is known at arrival, but generated length usually is not. As decoding proceeds, each active sequence claims more cache blocks. A batch that fit at tick 10 may overflow at tick 11 even when no new request arrives. The scheduler then needs a versioned rule for selecting a victim, preserving its logical progress, freeing physical blocks, and placing it back into a queue. This is the KV cache pressure point that makes LLM request preemption observable.
This is distinct from LLM admission control, which bounds entry before overload. It is also distinct from PagedAttention fragmentation, which studies how logical tokens occupy fixed-size blocks. The recovery ledger begins only when capacity has been exhausted for already-running work.
Treat one pressure event as the unit of evidence. Record the tick, active requests, capacity, occupied blocks, chosen victim, freed blocks, recovery mode, queue position, and eventual completion. The browser lab uses synthetic ticks and declared rates, not GPU timing. Its purpose is to make policy arithmetic falsifiable before an engineer spends a production rollout on it.
| Tick | Request | State | Blocks | Generated | Reason |
|---|---|---|---|---|---|
| 10 | R1 | running | 4 | 12 | admitted |
| 10 | R2 | running | 3 | 5 | admitted |
| 12 | R3 | preempted | 3 → 0 | 8 | capacity exceeded |
| 12–17 | R3 | queued | 0 | 8 | recovery work |
| 18 | R3 | running | 3 | 8 | progress preserved |
Reading rule: labels, markers, and the table carry every conclusion; color is supplementary.
Write the state that must survive eviction
A valid LLM request preemption receipt separates logical request state from physical cache residency. Logical state includes the request identifier, arrival tick, prompt tokens, generated tokens, remaining-token budget, priority class, age, and number of prior preemptions. Physical state includes occupied cache blocks, bytes per token, host-resident swap bytes, and whether restore or replay work remains.
Generated output is durable progress. If a request has already emitted 19 tokens, recovery cannot quietly reset that counter to zero. The user may have received those tokens, downstream code may have parsed them, and the remaining budget is defined relative to them. The lab therefore treats generated output as an invariant across every victim transition. A mutation test deliberately erases it and proves that the independent conservation oracle notices.
Prompt and output context need more careful labels. Recompute recovery frees modeled KV blocks and later replays the declared recoverable context: prompt tokens plus preserved generated tokens. A real engine may reuse prefix-cache state or apply implementation-specific optimizations, so this browser arithmetic is an upper-level policy model, not a prediction of kernel work. Swap recovery transfers the modeled resident bytes out and later back in; neither direction is described as free or zero-copy.
Shared prefixes, beam groups, speculative branches, and distributed cache ownership can make one request more than one allocation. Those are important production extensions, but inventing them in a small sandbox would obscure the core conservation rule. The bounded lab uses independent requests and flags this exclusion in every export. The receipt is useful precisely because it says what survives, what moves, and what the model does not contain.
Recompute turns memory pressure into prefill work
Under recomputation preemption, the scheduler discards the victim's modeled KV residency, returns its blocks to the capacity pool, and queues the request for recovery. When that request resumes, the engine must rebuild attention state for its prompt and preserved output context before it can generate another token. Memory pressure has become replay work.
Consider a synthetic victim with 24 prompt tokens and 8 generated tokens under an eight-token block size. Its context occupies ceil(32 ÷ 8) = 4 blocks. Evicting it frees four blocks. Recovery charges 32 recomputed tokens, not 24, because the eight emitted tokens still belong to the causal context. After replay, the generated-token counter remains eight and the remaining-token budget is unchanged. That is the conservation equation repeated in Figure 2, the semantic ledger, the lab, and its tests.
Current vLLM optimization documentation describes V1 preemption as recomputation and recommends reading cache size, concurrency, and scheduling configuration when warnings appear. This article does not instruct a reader to enable a removed CPU swap mode in current V1. Swap remains a systems comparison and a historical design option, not a current vLLM flag recommendation.
Recompute cost depends on model, prompt shape, prefix reuse, kernels, hardware, and batch interaction. The synthetic prefill-tokens-per-tick input only makes estimates comparable inside one receipt. LLM request preemption should therefore publish recomputed-token totals and queue delay separately from measured TTFT and TPOT. A modeled token count can justify the next experiment; it cannot borrow the latency of somebody else's accelerator.
Swap turns pressure into bounded transfer work
Swap recovery preserves a serialized representation of the victim's modeled KV state outside accelerator memory. In the sandbox, resident bytes equal context tokens multiplied by the declared KV bytes per token. A 32-token context at 512 modeled bytes per token occupies 16,384 bytes. A complete interruption transfers 16,384 bytes out and 16,384 bytes back, so the receipt reports 32,768 total transfer bytes.
Counting only swap-out makes the alternative look artificially cheap. Restore is required before decoding can continue, and host capacity is finite even when the transfer link is fast. The lab charges both directions and rejects a mutation that deletes restore bytes. It also records the declared host-transfer bytes per tick, because a byte total and a duration estimate answer different questions.
The PagedAttention paper describes all-or-nothing sequence eviction and compares swapping with recomputation in the system it evaluated. That design context is authoritative for the historical recovery model, not a transferable claim about today's hardware or current vLLM V1 behavior. A production decision needs versioned engine code, cache representation, host-memory budget, interconnect measurement, and workload traces.
Swap may appear attractive when recomputing a long context is expensive under declared rates. Recompute may appear attractive when transfer is slow or host memory is constrained. The adaptive-estimate mode calculates both synthetic estimates from the same victim state and chooses the lower one with a frozen tie rule. It records both estimates, so the choice can be challenged. LLM request preemption is auditable when the losing estimate remains visible instead of disappearing behind an “adaptive” label.
| Field | Recompute rail | Swap rail |
|---|---|---|
| Context | 24 prompt + 8 generated | 24 prompt + 8 generated |
| Freed blocks | ceil(32/8) = 4 | ceil(32/8) = 4 |
| Declared rate | 8 tokens/tick | 4096 bytes/tick |
| Charged work | 32 replay tokens | 16,384 out + 16,384 restore |
| Estimate | 4 ticks | 8 ticks |
| Boundary | synthetic arithmetic | synthetic arithmetic |
Reading rule: labels, markers, and the table carry every conclusion; color is supplementary.
Choose a victim without manufacturing starvation
A victim selector is a fairness policy disguised as a memory operation. Selecting the newest request can preserve work already invested in older sequences, but a burst of short arrivals may repeatedly displace the same long request after it resumes. Selecting the largest context frees more blocks, but can punish complex prompts. Selecting the lowest priority can make a background class wait forever unless the product contract defines an age escape hatch.
The lab freezes one transparent rule: among running requests, choose the lowest priority class first, then the newest arrival, then the lexicographically greatest stable identifier. Queue admission remains first-come, first-served within priority and arrival ties. This is not presented as universally optimal. It exists so two policy runs generate the same victim order and an independent oracle can reconstruct it without reading UI state.
Fairness guards observe rather than repair. A request is flagged when cumulative queue wait crosses the declared threshold or when its preemption count crosses the repeated-preemption threshold. The scheduler does not secretly promote it, because that would make the policy name incomplete. A production system may implement aging, reserved capacity, or class quotas, but those mechanisms must appear in the versioned rule and the replay.
Figure 3 uses non-color markers for the exact predicates: a triangle for excessive wait and a cross for repeated displacement. The alert is per request, not only a batch average. LLM request preemption can look healthy in aggregate while one request absorbs every recovery event, so the ledger retains victim history, age, queue ticks, completion tick, and starvation flags together.
Measure the policy with request-level evidence
A useful dashboard begins with counts but does not end there. Record pressure events, victims, freed blocks, replay tokens, swap-out bytes, restore bytes, queue re-entry, and completions. Join them to request arrival, prompt length, generated length, priority, first-token tick, completion tick, and final status. Only then can an engineer distinguish one harmless recovery from repeated displacement.
The current vLLM production metrics reference documents cache utilization, preemption, queue, prompt-token, TTFT, and request-level measures. The vLLM preemption metrics belong beside the request ledger, not in an isolated counter panel. Metric names and availability can change, so the deployment adapter should preserve engine version and collection configuration. A warning in logs is not a substitute for a trace joining the affected request to delay and completion.
Use distributions, not only means. Compare queue ticks and completion intervals by prompt bucket, output bucket, priority, tenant, and preemption count. Measure censored and failed requests explicitly; removing them can make the recovered population look faster. Chunked prefill scheduling changes how prefill shares compute with decode, so it belongs as a named workload condition rather than an explanation inferred after the fact.
The lab emits TTFT-like and TPOT-like synthetic intervals only to exercise the receipt shape. They are tick differences under a toy scheduler, not field latency. LLM request preemption rollouts should connect the same logical events to measured service metrics and LLM serving goodput gates. The release question is whether useful requests still meet declared promises, not whether a preemption counter became smaller in isolation.
| Request | Queue ticks | Preemptions | Completion | Alert |
|---|---|---|---|---|
| A | 4 | 0 | tick 28 | none |
| B | 13 | 1 | tick 41 | wait > 12 |
| C | 18 | 3 | tick 57 | wait > 12; preemptions > 2 |
| D | 7 | 2 | tick 45 | none; equality does not cross |
Reading rule: labels, markers, and the table carry every conclusion; color is supplementary.
Replay policy changes before production
Policy comparison requires an identical, versioned workload. Export sanitized arrivals, token counts, remaining budgets, and priority classes; freeze block size, capacity, modeled byte width, rate assumptions, work cap, tie rule, and fairness thresholds. Run recompute, swap, and adaptive-estimate against those same inputs. A changed policy should alter the event ledger hash while the normalized input hash remains stable.
Sensitivity analysis matters more than one winning row. Sweep KV capacity around the observed pressure boundary. Change synthetic prefill rate and transfer rate independently to locate the adaptive crossover. Preserve completion, victim order, starvation flags, and conservation checks for every run. If a small rate change flips the recommendation, the correct production conclusion is “measure the boundary,” not “adaptive solved it.”
Before rollout, define a gate: no erased output progress, no request beyond the allowed repeated-preemption threshold, bounded queue-delay regression, and acceptable goodput under a named traffic slice. Define rollback with the same event fields. Compare one engine version and one hardware shape at a time. Never transfer a throughput or latency result across models, GPUs, cache configurations, or prefix-sharing behavior without fresh measurement.
The downloadable lab accepts at most 64 requests and 20,000 ticks, validates unknown keys and impossible capacity before allocating its ledger, and uses no network, storage, or dynamic code. That narrowness is intentional. LLM request preemption becomes easier to review when the first artifact is small enough to replay byte for byte and hostile enough to reject ambiguous input.
Publish the preemption ledger
A durable preemption decision names engine version, policy version, victim rule, block model, capacity, rate assumptions, input hash, event hash, totals, conservation checks, fairness predicates, and limitations. It includes the losing adaptive estimate and both swap directions. It states that current vLLM V1 uses recomputation while swap is a historical systems comparison.
Review the receipt beside the production plan. Confirm that emitted output survives every transition, freed blocks equal the victim's modeled residency, recomputed context includes preserved output, and swap bytes include restore. Confirm that unknown fields, duplicate IDs, unsafe integers, impossible capacity, and over-cap runs fail before a large log is built.
Revisit this LLM request preemption guide on 2027-01-25, or sooner if vLLM changes behavior or metric names, a stable engine supersedes the cited recovery contract, search traffic expects a removed swap configuration, a source breaks, or field traces escape the sandbox assumptions. Refresh prose, figures, lab fixtures, and receipt hashes together.
The decision is not “recompute is always faster” or “swap preserves more work.” The decision is whether a named recovery rule conserves request progress, bounds resource work, exposes starvation, and survives a replay under the workload that matters.
Runnable local artifact — All ticks, rates, and bytes are declared synthetic arithmetic; this is not an LLM runtime, GPU benchmark, or current vLLM swap configuration guide.
Validate at most 64 requests before allocation, replay at most 20,000 ticks under a stable victim rule, preserve generated output, and export canonical input and event hashes.