Deterministic model-shape calculator · no runtime network

GQA vs MQA KV cache lab

Truth limit: This bounded lab calculates contiguous theoretical K/V payload for one sequence. It is not a GPU benchmark, quality evaluation, kernel-support oracle, serving-capacity model, or checkpoint-conversion tool.

Dimensions

Exact theoretical receipt

LayoutGQA
Sharing factor4×
Exact bytes134,217,728
Binary memory128 MiB

Query-to-K/V grouping

Dimension equation

What is and is not counted

The exact equation is layers × tokens × K/V heads × head dimension × bytes per element × two tensors. It excludes allocator metadata, padding, block fragmentation, sharding or replication, prefix reuse, temporary workspaces, model weights, and quality.