GQA vs MQA: Choose the Right KV Cache
Map query heads to shared K/V heads, calculate exact cache payload, and choose a point that clears capacity, quality, and kernel gates.
GQA vs MQA changes how query heads share key/value heads, which directly changes the theoretical KV cache payload during autoregressive decoding. This guide maps the head wiring, calculates exact bytes, and separates architecture math from the quality, kernel, and runtime evidence required for a real model decision.
GQA vs MQA begins with the cache budget
GQA vs MQA is a memory-layout decision before it is a quality claim. During autoregressive decoding, every layer retains one key vector and one value vector per past token and per key/value head. Multi-query attention (MQA) shares one key/value head across all query heads; grouped-query attention (GQA) shares several, with each key/value head serving an equal query-head group. Both reduce the KV cache relative to ordinary multi-head attention (MHA).
The exact theoretical byte ledger is straightforward: layers × cached tokens × KV heads × head dimension × bytes per element × 2 for keys and values. Query-head count does not directly multiply cache bytes. It determines the sharing factor and must divide evenly by the key/value-head count for the simple grouped mapping used here.
The original multi-query attention paper describes sharing keys and values across attention heads to reduce incremental decoding memory bandwidth. That establishes the mechanism, not a universal claim that every existing checkpoint can be converted without training or that one layout wins on every runtime.
This guide gives you a dimensionally correct calculator, an explicit head map, and a decision frontier. Use those artifacts to choose what to train or serve under a measured cache constraint, while keeping quality evaluation, kernel support, tensor parallelism, and checkpoint compatibility as separate gates.
| Layout | Query heads | K/V heads | Sharing |
|---|---|---|---|
| MHA | 8 | 8 | 1 query per K/V |
| GQA | 8 | 2 | 4 queries per K/V |
| MQA | 8 | 1 | 8 queries per K/V |
Reading rule: labels, values, patterns, and structure carry the conclusion; color is supplementary.
Draw the query-to-KV head wiring
Start with an example that can be audited by sight: eight query heads. MHA pairs them with eight key heads and eight value heads. GQA with two KV heads assigns query heads 0–3 to KV head 0 and query heads 4–7 to KV head 1. MQA assigns all eight query heads to the single KV head. The same assignment applies to keys and values; unequal key- and value-head counts are outside this contract.
For a simple contiguous grouping, the group size is queryHeads / kvHeads and the mapping is floor(queryHead / groupSize). Reject a configuration when queryHeads mod kvHeads is nonzero. A silent rounding rule would produce uneven groups and a model shape different from the one the byte calculator claims to represent.
That wiring makes the cache tradeoff concrete. Moving from eight KV heads to two divides theoretical KV bytes by four; moving from eight to one divides them by eight. It does not reduce the number of query heads or imply that attention computation, output projection, latency, or model quality improves by the same factor.
The related multi-head latent attention KV-cache guide studies a different compression contract. Latent attention changes what is cached and reconstructed; grouped sharing changes how query heads reuse explicit key/value heads. Compare their ledgers, but do not treat them as equivalent parameter substitutions.
Calculate KV cache bytes without hidden units
Keep every factor named and every unit visible. For 32 layers, 4,096 cached tokens, 128 dimensions per head, FP16 storage at two bytes per element, and two KV heads, the ledger is 32 × 4,096 × 2 × 128 × 2 × 2 = 134,217,728 bytes. That is 128 MiB when divided by 1,048,576. MQA with one KV head is 64 MiB; eight KV heads are 512 MiB under the same assumptions.
Those numbers cover stored K and V tensors for one sequence. They exclude allocator metadata, block rounding, fragmentation, temporary workspaces, attention scores, activations, model weights, prefix sharing, speculative branches, and runtime-specific layouts. Batch residency is not always “single-sequence bytes × requests,” because sequence lengths differ and engines may share or page blocks.
The companion KV-cache quantization error guide examines the separate bytes-per-element lever. Reducing KV heads changes the attention architecture; quantizing the cache changes representation precision. You can combine them, but the resulting quality and kernel behavior must be evaluated as a combined configuration rather than adding two isolated claims.
Use bytes first and label binary units correctly. If a product requirement is expressed in GB, state whether it means 10⁹ bytes or GiB. Small naming shortcuts become large capacity errors across many layers, tokens, and concurrent sequences.
Separate architecture evidence from runtime evidence
A theoretical ledger answers “how many payload bytes does this tensor shape require?” It does not answer “how many requests fit on this GPU?” A serving runtime allocates blocks, reserves memory, selects kernels, shards heads, and may duplicate or communicate state across devices. Measure actual reserved and resident bytes on the target runtime after the shape passes the paper calculation.
The GQA paper describes grouped-query attention as an interpolation between multi-head and multi-query attention and studies uptraining multi-head checkpoints. Read its empirical results within the reported models, data, and procedure. They do not license arbitrary checkpoint surgery, nor do they guarantee that a new model will keep its quality at a chosen group count.
Kernel support is another gate. PyTorch’s scaled dot product attention reference documents current grouped-query constraints and an enable_gqa interface. Confirm the installed version, backend, device, head divisibility, and key/value equality rather than copying an example into a production assumption.
For GQA vs MQA, publish two receipts: the architecture receipt names Q heads, KV heads, head dimension, layers, element width, and training provenance; the runtime receipt names engine, version, device, parallelism, block size, allocated bytes, throughput, and latency. Agreement between them is evidence. A gap is an investigation, not an invitation to rename overhead as cache payload.
- Exact formula
- layers × cached tokens × K/V heads × head dimension × bytes per element × two tensors.
- Declared example
- 32 × 4,096 × 2 × 128 × 2 × 2 = 134,217,728 bytes = 128 MiB.
- Boundary
- Allocator, paging, sharding, prefixes, workspaces, and metadata are excluded.
Reading rule: labels, values, patterns, and structure carry the conclusion; color is supplementary.
Place GQA vs MQA on a decision frontier
Treat KV-head count as a discrete frontier, not a winner-take-all benchmark. The left edge, MQA, minimizes theoretical cache payload for fixed dimensions. Increasing KV heads spends memory to give query groups more distinct key/value representations. MHA sits at the right edge when KV heads equal query heads. Candidate points between them are GQA configurations only when divisibility and runtime support hold.
Begin with the maximum resident context and concurrency target. Convert that capacity into a per-sequence payload ceiling after reserving model weights, runtime overhead, and a safety margin. Eliminate head layouts whose theoretical bytes already exceed the ceiling. Then benchmark the surviving layouts on the actual engine and evaluate the model on tasks sensitive to long context, retrieval, multilingual behavior, and generation stability.
The PagedAttention fragmentation article helps translate payload bytes into blocked runtime residency. Paging can reduce waste from reserved-but-unused sequence capacity, yet it cannot erase the K/V payload each live token requires. Put the payload ledger and fragmentation measurement on adjacent rows.
A practical frontier table therefore has columns for KV heads, sharing factor, theoretical bytes per sequence, measured resident bytes, tokens per second, time to first token, inter-token latency, evaluation deltas, and operational status. Select the smallest-memory point that clears every quality and compatibility floor, not the point with the most flattering isolated metric.
Test quality with a matched model contract
Memory math is deterministic; quality is empirical. Compare GQA vs MQA only when model size, tokenizer, training or uptraining procedure, data mixture, context distribution, decoding settings, and evaluation harness are documented. If those facts differ, attribute the result to the complete model recipe rather than to one head-layout label.
For a model being trained, choose several divisible KV-head candidates early enough that training can adapt. For an existing checkpoint, use an officially supported architecture or a documented uptraining procedure. Repeating or averaging tensors to force a new count may create a runnable file, but it is not evidence of retained model behavior. This article deliberately makes no arbitrary-checkpoint conversion recipe.
Evaluate both aggregate and sliced behavior. Long-context retrieval, repeated entities, multilingual prompts, code completion, tool-call argument stability, and multi-turn instruction retention can react differently. Freeze prompts and decoding seeds where possible, report confidence intervals or paired outcomes, and inspect regressions rather than hiding them in a single average.
The cache calculator stays useful here because it gives each candidate an exact memory coordinate. Pair that coordinate with quality thresholds defined before evaluation. When no candidate clears both capacity and quality, change the product constraint, train a different architecture, alter precision, or reduce supported context; do not manufacture a passing conclusion by moving the threshold after seeing results.
- Calculate exact payload for each divisible K/V-head count.
- Eliminate candidates above the cache budget.
- Evaluate matched model recipes against frozen quality floors.
- Profile engine support, measured residency, and parallel topology.
- Choose the lowest-memory candidate that passes every gate.
Reading rule: labels, values, patterns, and structure carry the conclusion; color is supplementary.
Model concurrency, prefixes, and eviction separately
Per-sequence cache bytes are an input to capacity planning, not the whole scheduler. Real traffic mixes prompt lengths, decode lengths, cancellations, shared prefixes, and idle gaps. Build a trace model that accounts for live token residency over time and then apply the chosen layout’s bytes per token. Keep predicted and observed occupancy in the same unit.
Prefix reuse can avoid duplicate prefill computation and storage when identity rules permit it. Eviction decides which reusable inactive blocks remain under pressure. The KV-cache eviction guide covers active-reference safety, ancestry, tenant scope, and avoided-prefill value. GQA vs MQA changes the byte cost of each cached token; it does not choose victims or make cross-tenant reuse safe.
Include the worst credible mixture: maximum supported context, a burst of long prompts, minimal prefix overlap, and decode tails that keep cache resident. Also include the common mixture so the service is not designed only for an unlikely corner. Use admission control when the worst case cannot fit rather than relying on an out-of-memory failure as policy.
Shard topology matters. When KV-head count interacts poorly with tensor-parallel width, a theoretically compact candidate can require replication or communication that changes latency and memory. Record the mapping explicitly. This is why the final decision belongs to a model-and-runtime pair, even though the core byte formula remains portable and exact.
Ship the calculator and the decision receipt
The local GQA vs MQA lab accepts bounded integers for query heads, KV heads, head dimension, layers, tokens, and bytes per element. It rejects nondivisible head layouts, requires one equal K/V-head count, computes exact integer bytes, reports binary MiB and GiB, and renders the tiny reference grouping for every query head. JSON and CSV exports preserve inputs, formulas, outputs, and validation status.
The GQA vs MQA lab is not a GPU benchmark, model evaluator, kernel compatibility oracle, or conversion tool. Its byte total describes contiguous K/V payload for one sequence under the declared dimensions. Runtime overhead, paging, sharding, quantization metadata, prefix reuse, and allocator behavior require separate measurements. The fixed 8/2/1 preset exists to make the grouping logic inspectable, not to recommend two KV heads.
For release, keep one signed-off row per candidate: checkpoint identity and provenance; Q and KV heads; head dimension; layer count; cache dtype; maximum context; theoretical payload; measured memory; engine and kernel versions; parallel topology; performance slices; quality gates; and rollback target. Recalculate when any dimension changes.
That is the honest choice: use the exact ledger to eliminate impossible layouts, use matched evaluations to protect behavior, use runtime profiling to confirm residency and speed, and choose the surviving frontier point. The arithmetic is small. The engineering discipline is refusing to make it claim more than it measures.
Runnable local artifact — The lab calculates contiguous theoretical K/V payload for one declared sequence; it is not a GPU benchmark, quality evaluation, kernel-support oracle, serving-capacity model, or checkpoint-conversion tool.
Validate equal K/V heads and divisibility, map every query head to a group, compute exact integer bytes and sharing factor, and export the full dimensional receipt.