Deterministic model-shape calculator · no runtime network
GQA vs MQA KV cache lab
Truth limit: This bounded lab calculates contiguous theoretical K/V payload for one sequence. It is not a GPU benchmark, quality evaluation, kernel-support oracle, serving-capacity model, or checkpoint-conversion tool.
Dimensions
Exact theoretical receipt
LayoutGQA
Sharing factor4×
Exact bytes134,217,728
Binary memory128 MiB
Query-to-K/V grouping
Dimension equation
What is and is not counted
The exact equation is layers × tokens × K/V heads × head dimension × bytes per element × two tensors. It excludes allocator metadata, padding, block fragmentation, sharding or replication, prefix reuse, temporary workspaces, model weights, and quality.