Home›Journal›This post

EAGLE vs N-Gram Speculative Decoding

Compare learned draft trees with prompt lookup through a paired workload receipt that separates proposal, verification, queueing, quality, and operating cost.

JP
JP Casabianca
AI Engineer and Product Designer · full-stack delivery · Bogotá

EAGLE vs n-gram speculative decoding is a workload decision, not a leaderboard decision. This guide builds a paired trace receipt that reveals when repeated prompt text rewards lookup, when a learned draft tree earns its cost, and when target verification erases both advantages.

EAGLE vs n-gram speculative decoding starts with the trace

The useful question is not which technique won a published benchmark. It is which complete serving path lowers time per emitted token for one target model, engine revision, sampling policy, workload mix, and concurrency cohort without changing output quality. EAGLE vs n-gram speculative decoding becomes answerable only after those conditions are attached to every row.

Both proposers remain subordinate to target verification. The n-gram path looks for a continuation already visible in prompt or generated text. EAGLE-3 predicts a tree from learned features. In either case, the target model decides what may be emitted, so proposed tokens, accepted draft tokens, and authoritative emitted tokens must stay separate.

Start with a paired trace receipt: target and draft revisions, engine and configuration digests, hardware label, sampling digest, workload hash, concurrency cohort, warmup policy, request IDs, output policy, and units. The current vLLM speculative-decoding example exposes distinct configuration surfaces for EAGLE/EAGLE-3 and n-gram lookup; that is implementation evidence, not a promise that either wins your traffic. For the distribution-preserving mechanism itself, begin with speculative decoding foundations.

Two proposers, one verifierA dashed cyan lookup rail and branching coral learned-draft rail converge on a white target-verification gate. Solid tokens pass; crossed branches are rejected.PROPOSAL IS ADVISORY · TARGET EMISSION IS AUTHORITATIVE1 · LOOKUP123452 · LEARNED DRAFT3 · TARGETVERIFY4 · EMIT
Two proposers, one verifier. Lookup and learned candidates remain proposals until one matched target verifies them.
  1. The n-gram rail searches repeated prompt or output spans without learned weights.
  2. EAGLE-3 grows a learned token tree from target features.
  3. The same target model verifies either proposal path.
  4. Only verified authoritative tokens are emitted; rejected branches never count as output.

N-gram lookup earns tokens from repetition

An n-gram proposer is a source-visible copier. It searches bounded token spans already present in the request context and proposes the continuation after a match. There are no learned draft weights to load, distribute, or keep compatible with a target checkpoint. That small operational surface is its first advantage.

Its second advantage is conditional: repeated structure. Repository edits often echo imports, braces, identifiers, or a nearby function. Tool payloads repeat JSON keys and schema phrases. Retrieval answers may quote prompt material. In those cohorts, a lookup can produce a useful run while doing little proposer work. A short match window may be cheap but ambiguous; a long one may be selective but rare. Record both the window and maximum draft length.

Novel reasoning, fresh prose, and rapidly diverging code erase that leverage. The n-gram proposer can offer zero accepted tokens without being broken; the workload simply lacks reusable continuation. Position matters too: copied context near the prompt may help early, while later tokens drift away. EAGLE vs n-gram speculative decoding therefore needs repetition buckets and output-position slices, not one blended acceptance average. The frozen teaching matrix keeps copied code, structured JSON, conversational text, and novel reasoning separate so a favorable cohort cannot conceal an irrelevant one.

EAGLE-3 pays for a learned draft tree

EAGLE-3 takes a different bet: learned prediction can draft useful tokens even when the context contains no exact reusable span. The EAGLE-3 paper describes direct token prediction and multi-layer feature fusion, then evaluates that design under specific models, hardware, tasks, and serving conditions. Those results establish a mechanism and bounded experiments; they do not transfer as a multiplier to this article's synthetic fixture.

The learned path introduces real costs. A compatible draft checkpoint has to exist for the target. Weights consume memory and need release lineage. Feature extraction and tree growth take time. The engine must support the proposal and verification topology, and operators need metrics that distinguish draft work from target work. A target upgrade can invalidate the draft relationship even when an n-gram proposer still runs.

That price may be worthwhile in low-repetition cohorts where useful branches survive verification. Tree verification can amortize a target pass across candidates, yet a larger tree is not automatically better: rejected branches still consumed proposal and verification work. Record the draft revision, tree width, draft length, memory delta, and verifier share. EAGLE vs n-gram speculative decoding is a comparison of these full paths, not learned intelligence against a supposedly free string trick.

Measure proposal, verification, and acceptance separately

Acceptance length counts accepted draft tokens per proposal step. Acceptance ratio divides accepted draft tokens by proposed tokens. Neither is latency. A method may accept longer runs while spending enough proposer or target-verification time to lose on measured time per authoritative emitted token. Use weighted totals across rows; averaging row ratios lets tiny requests distort the result.

For each request, sum proposer, verifier, other execution, and queue milliseconds. Divide the execution total by authoritative emitted tokens only when the denominator is positive. Report queue separately as well, because concurrency can make scheduler delay dominate the serving path. Keep TTFT and inter-token latency beside throughput metrics rather than compressing them into one score.

The frozen fixture has 48 paired copied-code requests, 44 paired novel-reasoning requests, and 52 paired structured-JSON requests, evenly split between the declared position bands. In the position-1 copied-code slice, lookup records 9 accepted of 12 proposed tokens at 4.8 ms per emitted token, while EAGLE-3 records 8 of 14 at 5.2 ms. Novel reasoning reverses the measured order: lookup accepts 1 of 8 and records 8.1 ms, while EAGLE-3 accepts 7 of 12 and records 6.4 ms. Structured JSON differs by 0.1 ms in both bands, inside the 0.25 ms review threshold, and is labeled inconclusive. These are synthetic arithmetic values, never GPU measurements.

The lab uses nearest-rank percentiles: sort request values and choose rank ceil(p × n), clamped to the available range. It also reports verifier share as total verifier time divided by total execution time. Those definitions make EAGLE vs n-gram speculative decoding recomputable from exported rows.

Acceptance-by-position cohort mapA patterned three-row matrix shows lookup leading for repeated code, EAGLE-3 leading for novel reasoning, and structured JSON inside an inconclusive band across two frozen concurrency and output-position slices.SYNTHETIC COHORTS · READ PATTERN + LABEL, NOT COLOR ALONEC1 · POSITION 1C24 · POSITION 241 · Copied codelookup 4.8n=24 / 24lookup 5.7n=24 / 242 · Novel reasoningEAGLE-3 6.4n=22 / 22EAGLE-3 7.8n=22 / 223 · Structured JSONinconclusiven=26 / 26inconclusiven=26 / 26
Acceptance-by-position cohort map. One fixture deliberately contains a lookup cohort, a learned-draft cohort, and a no-winner cohort.
Frozen synthetic cohort values
CohortC1 · position 1C24 · position 24SamplesWinning-method verifier share
Copied coden-gram 4.8 vs EAGLE-3 5.2 ms/tokenn-gram 5.7 vs EAGLE-3 6.0 ms/token48 pairedn-gram 31.25% / 44%
Novel reasoningn-gram 8.1 vs EAGLE-3 6.4 ms/tokenn-gram 8.2 vs EAGLE-3 7.8 ms/token44 pairedEAGLE-3 38% / 52%
Structured JSONn-gram 6.0 vs EAGLE-3 6.1 ms/tokenn-gram 7.0 vs EAGLE-3 7.1 ms/token52 pairedn-gram 35% / 49%; both inconclusive

Reading rule: labels, patterns, markers, and the semantic content carry every conclusion; color is supplementary.

Slice by repetition, task, and concurrency

A single dashboard tile destroys the evidence needed for a proposer choice. Preserve task family, repetition bucket, output-position band, output length, and concurrency cohort on every trace. Then require minimum paired sample counts before interpreting a delta. The cohort matrix is a map of where to investigate, not permission to cherry-pick a winner.

Concurrency changes the question. At batch size one, draft overhead and host coordination may be conspicuous. Under higher concurrency, target verification competes for batch capacity and queue time can swamp a local token saving. Prefix reuse may reduce target work in one engine configuration but interact differently with a draft path. Compare engines only after their scheduling contract is matched; vLLM vs SGLang for agent workloads is the wider engine decision, while this receipt holds the engine fixed.

Use at least p50 and p95 per-request time per emitted token, plus total weighted values. Tail behavior catches a proposer that looks economical on short outputs but creates expensive verification bursts. Study low and high concurrency independently, then consult continuous batching for the scheduler mechanics. EAGLE vs n-gram speculative decoding can legitimately yield different choices for copied code at concurrency 1 and mixed chat at concurrency 32.

Reject incomparable or quality-divergent runs

Eligibility comes before arithmetic. Reject the pair when target model or revision, engine/config digest, hardware, sampling digest, workload hash, concurrency cohort, measurement units, warmup policy, request IDs, or quality policy differ. A faster run against a smaller workload is not an approximate comparison; it is a different experiment.

For deterministic decoding, compare normalized output digests. For stochastic evaluation, pin seeds where supported and declare an approved quality evaluator, tolerance, and sample plan. Do not call unequal text a serving win merely because it ended sooner. Reasoning budgets must also match; test-time compute stop rules covers that separate policy boundary.

Warm every serving path with the same rule, retain failures, and pair by request ID. Report zero emitted-token rows as undefined for time-per-token rather than dividing by zero. Never import a headline speedup from a paper into a local receipt. A recent workload study, Speculative Decoding: Performance or Illusion?, reports sensitivity to workload, batch conditions, acceptance length, and target-verification cost under its own setups. EAGLE vs n-gram speculative decoding uses that as a reason to measure, not as a substitute for measurement.

Net-cost waterfallThree horizontal rails decompose synthetic queue, proposal, target verification, and other emitted-token time for n-gram, EAGLE-3, and no speculation.ONE SYNTHETIC FIXTURE · MILLISECONDS PER EMITTED TOKENn-gram7.48 msEAGLE-37.32 msnone7.00 msgray queue · hatched proposer · coral verifier · cyan other execution
Net-cost waterfall. The lowest total depends on every cost component; accepted tokens alone cannot select a proposer.
Fixture
144 paired synthetic requests: 48 copied-code, 44 novel-reasoning, and 52 structured-JSON; not a GPU benchmark.
n-gram
0.90 queue + 0.71 proposer + 3.10 verifier + 2.77 other = 7.48 ms per emitted token end to end.
EAGLE-3
0.90 queue + 1.08 proposer + 2.88 verifier + 2.46 other = 7.32 ms per emitted token end to end.
No speculation
0.90 queue + 0 proposer + 3.90 verifier + 2.20 other = 7.00 ms per emitted token end to end.
Decision
The no-speculation control is lower in this complete frozen fixture; cohort slices still show conditional proposer behavior.

Price the operational surface

Performance is only one ledger. EAGLE-3 needs checkpoint availability, distribution, memory reservation, compatibility checks, and rollback coordination with the target. N-gram lookup needs bounded search settings and observability but no learned artifact. Either may be unavailable or immature in the chosen engine version, which is a valid decision outcome.

Attach deployment facts to the receipt: added GPU memory, image or artifact size, startup time, configuration ownership, metric coverage, and rollback steps. Include a no-speculation control using the same target and workload. If both proposers regress against that control, the decision is not “pick the less bad accelerator”; it is “disable speculation for this cohort.”

Operational cost can reverse a narrow latency win. A two-percent synthetic improvement does not justify a fragile checkpoint pipeline by default, while a simple lookup improvement on a dominant copied-code cohort might. Conversely, a stable learned draft with a meaningful tail reduction can earn its surface. State the threshold before looking at results. The lab uses an explicit review band and returns inconclusive inside it, because EAGLE vs n-gram speculative decoding should preserve a no-winner state.

Publish the proposer decision receipt

A useful decision table has one row per eligible cohort and columns for sample count, proposed, accepted, emitted, weighted acceptance ratio, p50/p95 time per emitted token, verifier share, queue share, quality result, control delta, and operational note. Add rejected cohorts beneath it with the exact mismatch reason. Silence is not evidence.

Publish target/draft revisions, engine/config and workload hashes, sampling, hardware, units, warmup, draft settings, collection window, and formulas. Label the frozen illustration as synthetic. Store raw sanitized rows beside the aggregate receipt so a reviewer can reproduce every number. This creates an update path when traffic or software changes.

The decision may be lookup for repeated code, EAGLE-3 for low-repetition reasoning, neither under saturated concurrency, or inconclusive while more data is collected. That conditional result is more useful than a universal badge. EAGLE vs n-gram speculative decoding becomes an engineering choice only when the cost of proposing, verifying, queueing, operating, and preserving output quality stays attached to the same trace.

Runnable local artifact — The frozen values are synthetic teaching data; the lab is neither an inference engine nor an EAGLE simulator and names no universal winner.

Plain text1 line
Validate at most 512 requests and 8,192 proposal steps, reject mismatched comparison keys, compute weighted acceptance and nearest-rank percentiles, and export one canonical receipt.