Home›Journal›This post

MXFP4 Quantization: Audit the 4.25-Bit Blocks

Decode one MXFP4 block, verify its 4.25-bit arithmetic, enumerate E2M1 values, expose outlier damage, and publish an exact logical receipt.

JP
JP Casabianca
AI Engineer and Product Designer · full-stack delivery · Bogotá

Four bits per element is not four bits per stored value. In MXFP4, 32 E2M1 elements share one eight-bit E8M0 scale, so the logical block is 136 bits—or 4.25 bits per value—before container metadata, alignment, and mixed-precision tensors enter the picture. This MXFP4 quantization audit turns that block into a deterministic receipt instead of a marketing label.

MXFP4 quantization is a block, not a nibble

MXFP4 quantization starts with four-bit elements, but the deployable claim belongs to a block. Thirty-two FP4 E2M1 values share one eight-bit E8M0 scale. The logical accounting is therefore 32 × 4 + 8 = 136 bits, and 136 ÷ 32 = 4.25 bits per weight. Leaving out the scale byte turns a block format into an inaccurate four-bit slogan.

That arithmetic is deliberately narrow. It does not include tensor names, shapes, offsets, alignment, container metadata, padding, checksums, unquantized tensors, or a runtime's expanded working set. The OCP Microscaling Formats specification defines logical encodings and minimally supported conversions; it does not prescribe one universal physical memory layout.

The useful unit of inspection is one complete block receipt: 32 source numbers, one declared scale policy, one scale byte, 32 nibbles, 32 reconstructed numbers, and error counts. That receipt lets an engineer reproduce the claim without borrowing a vendor kernel or trusting a filename.

Use the 4.25 bits per weight figure as logical payload arithmetic. For a checkpoint or serving decision, measure the actual artifact and process separately. LLM quantization quality budgets begin where this block audit ends: model-level quality, memory, and latency evidence.

136-bit block anatomyOne eight-bit scale spans thirty-two four-bit element cells grouped into sixteen lab bytes.ONE SCALE · THIRTY-TWO NIBBLESE8M0 SCALE · 8 BITS0123456789101112131415161718192021222324252627282930318 + 32 × 4 = 136136 ÷ 32 = 4.25 logical bits/value
MXFP4 block accounting includes one shared scale byte and all thirty-two nibbles.
Logical block accounting
FieldValueMeaning
Scale width w8 bitsone E8M0 scale
Element width d4 bitsone E2M1 nibble
Block size k32shared-scale values
Formula8 + 32×4136 logical bits
Per value136/324.25 logical bits
Reading rule
Read each row across its named columns; the text carries the diagram's exact values.
  • Color and position reinforce the comparison but never replace its labels.

Read the FP4 E2M1 value lattice

FP4 E2M1 has one sign bit, two exponent bits, and one explicit mantissa bit. In the finite MXFP4 interpretation used here, its nonnegative magnitude lattice is 0, 0.5, 1, 1.5, 2, 3, 4, and 6. The sign bit mirrors those magnitudes, including positive and negative zero, across all 16 nibbles.

The 0.5 point is subnormal. It matters because the midpoint between zero and 0.5 is 0.25: values below that midpoint round to zero, while an exact halfway case needs a declared tie rule. The reference lab uses round-to-nearest, ties-to-even over the nibble lattice. An exact tie chooses the candidate whose least-significant mantissa bit is even; it does not simply round away from zero.

After dividing a source value by the shared scale, magnitudes beyond 6 cannot be represented. The lab clamps them to signed 6 and increments saturation. Tiny values may become signed zero and increment underflow when the nonzero source reconstructs as zero. Those names describe encoding outcomes, not model quality.

Enumerating the lattice is more reliable than reimplementing a compact bit trick from memory. The audit exposes every nibble, sign, magnitude, and reconstructed value, so an independent decoder can compare all 16 cases. FP8 calibration range covers a different format and calibration problem; its range intuition should not be substituted for this FP4 E2M1 table.

Read the shared E8M0 scale

The scale is an unsigned eight-bit exponent code for a power of two. Decode an ordinary scale byte e as 2^(e−127); the all-ones code is reserved for NaN in the logical E8M0 format. A block therefore moves the entire E2M1 lattice up or down together rather than assigning 32 independent scales.

That sharing is the central compression trade-off. A well-scaled block can use several lattice levels. A single large value can demand a larger exponent, widen the dead zone for the other 31 values, and reconstruct many of them as zero. Conversely, a scale that is too small saturates the largest magnitudes. MXFP4 quantization does not erase this tension; it makes the chosen E8M0 block scale inspectable.

The lab offers two explicit paths. Its OCP §6.3 max-based path computes the nonzero block exponent as floor(log2(maximum absolute value ÷ 4)), then clamps that exponent to the valid E8M0 range from −127 through 127. This is not the same as choosing a scale that always keeps the normalized maximum at or below 6: the specified rule can deliberately leave a top value above the finite lattice. Manual mode accepts the same full exponent range so reviewers can reproduce a known fixture or probe an alternative. The all-zero block uses a named lab policy with exponent zero because there is no unique data-driven scale to infer.

These are scalar audit rules. A production converter may make additional choices permitted by its versioned contract. Record the policy identifier and implementation version instead of presenting one scale-selection procedure as the entire specification.

Encode one MXFP4 block step by step

Take a synthetic block whose largest magnitude is 13. The OCP §6.3 max-based rule gives floor(log2(13 ÷ 4)) = 1, so the scale is 2. That choice leaves 13 ÷ 2 = 6.5 above the largest finite lattice magnitude 6; saturation is an observable consequence of the policy, not a reason to silently select the next scale. Each source value is divided by 2, rounded to the nearest E2M1 value with ties-to-even, encoded as a nibble, then multiplied by 2 to reconstruct. That makes MXFP4 quantization inspectable at the exact block where the shared scale acts.

For a source of 3, the normalized value is 1.5, an exact lattice point, so reconstruction is 3. A source of 5 normalizes to 2.5, exactly between 2 and 3; the even tie rule resolves through the encoded mantissa parity. A source of 0.3 normalizes to 0.15 and underflows to zero. The 13 source normalizes to 6.5 and saturates at 6, reconstructing 12.

Nibble packing is logical and explicit: values 0 and 1 occupy the low and high half of the first lab byte, values 2 and 3 the next, and so on. The lab prepends the scale code only for its own deterministic serialization. That byte order is labeled “lab serialization,” never “checkpoint layout.”

The canonical receipt preserves source strings after numeric validation, normalized values, nibbles, packed lab bytes, reconstructions, and metrics. Replaying the same block and policy must produce byte-for-byte identical JSON. The current Torch-TensorRT quantization guide is useful as one versioned packed E2M1/E8M0 integration path, not evidence that every runtime stores the same bytes.

E2M1 lattice and scale ladderThe eight finite magnitudes appear at three scale steps with distinct outcome shapes.ROUND ON THE LATTICE · THEN APPLY THE SCALE00.511.52346× 1× 2× 4circle: zero · amber: subnormal · mint: exact or roundedcoral edge: saturation
The finite E2M1 magnitude lattice is scaled by one power of two for the whole block.
E2M1 magnitude encodings
MagnitudePositive nibbleNegative nibbleRole
000001000signed zero
0.500011001subnormal
100101010finite
1.500111011finite
201001100finite
301011101finite
401101110finite
601111111finite maximum
Reading rule
Read each row across its named columns; the text carries the diagram's exact values.
  • Color and position reinforce the comparison but never replace its labels.

Expose outlier and dead-zone damage

A contact sheet becomes useful when it changes one controlled variable. The exact fixture repeats −2, −1.5, −1, −0.5, 0, 0.5, 1, 1.5, and 2 until it has 31 values, then appends either 5 or 20. The logical payload remains 136 bits in both fixtures, but the shared scale rises from 1 to 4, coarsening every reconstructed step.

Under the OCP section 6.3 max rule, the baseline exponent is 0. Its final 5 lands exactly between the E2M1 values 4 and 6, so ties-to-even reconstructs 4; the fixture has 3 reconstructed zeros, maximum absolute error 1, and RMSE 0.176777. Replacing only that value with 20 produces exponent 2: 20 ÷ 4 again lands at 5, reconstructs as 16, and the full block has 17 zeros, maximum absolute error 4, and RMSE 0.910014. The figure and semantic table are generated from this same independent fixture oracle.

Report zero, underflow, and saturation counts beside maximum absolute error and root mean squared error. Maximum error surfaces the worst reconstruction. RMSE summarizes energy across the block but can hide a concentrated failure, so neither metric should replace the per-element table. Relative error is shown only for nonzero source values; dividing by zero would manufacture a meaningless infinity.

The outlier experiment is a block-level diagnostic. It does not predict perplexity, downstream accuracy, or the behavior of a full model. It does reveal whether an MXFP4 quantization scale policy spends most of its lattice on one value while turning quieter values into a dead zone. That is original value a checkpoint size cannot provide.

MXFP4 quantization receipts should keep the unchanged bit count visible when error changes. Otherwise a viewer may mistake the experiment for a compression comparison. The visual uses numeric labels and shape codes for exact, rounded, underflowed, and saturated cells so the conclusion does not depend on mint, cobalt, or coral alone.

Separate logical bits from checkpoint bytes

A logical 4.25 bits per weight statement is not a file-size formula. Real containers need tensor metadata, alignment, and offsets. A model may leave embeddings, normalization parameters, routing state, or output heads in wider formats. Compression tools may add indexes or pad packed blocks. Serving systems may materialize dequantized tiles, caches, and workspace buffers.

The OpenAI gpt-oss model card makes a bounded claim: its mixture-of-experts weights are quantized to MXFP4 at 4.25 bits per parameter. That sentence should not be expanded into “every parameter uses 4.25 bits” or “the checkpoint is exactly parameters × 4.25 bits.” Measure the downloaded files and inspect tensor dtypes to answer those questions.

Safetensors versus GGUF compares container and runtime trade-offs after the logical format is understood. The block lab intentionally avoids naming a universal physical order. Its serialization exists so tests can hash one deterministic teaching artifact, not so a loader can assume the same representation.

Keep three numbers separate in reviews: logical MX payload bits, actual checkpoint bytes, and peak runtime memory. Each answers a different question, and only the first is derived by 8 + 32 × 4.

Outlier pressure contact sheetTwo exact thirty-two-value fixtures keep 136 bits while one outlier changes scale zeros and error.SAME 31-VALUE PREFIX · CHANGE ONLY THE LAST VALUEBASELINE · MAX 5OUTLIER · MAX 20scale 1 · zeros 3 · 5 → 4scale 4 · zeros 17 · 20 → 16max error 1 · RMSE 0.176777max error 4 · RMSE 0.910014logical payload · 136 bitslogical payload · 136 bitsOCP §6.3 exponent = floor(log2(max / 4))mint exact · amber zeroed · cobalt rounded · coral largest error
An exact synthetic outlier raises the shared scale and reconstruction error without changing logical storage.
Executable synthetic contact sheet
FixtureExact constructionScale exponentScaleLast valueZerosMax errorRMSEBits
Baselinerepeat [-2,-1.5,-1,-0.5,0,0.5,1,1.5,2] to 31 values; final 5015 → 4310.176777136
One outlierrepeat [-2,-1.5,-1,-0.5,0,0.5,1,1.5,2] to 31 values; final 202420 → 161740.910014136
Reading rule
Read each row across its named columns; the text carries the diagram's exact values.
  • Color and position reinforce the comparison but never replace its labels.

Validate the model and runtime separately

Block parity is the first gate, not the release verdict. Decode known nibbles independently, compare packed receipts, and run boundary fixtures for signed zero, subnormal 0.5, exact ties, saturation, exponent limits, and the all-zero policy. Then inspect tensor coverage so the percentage of values actually using MXFP4 quantization is explicit.

Model evaluation belongs on representative tasks with a stated baseline, dataset revision, decoding setup, and acceptance budget. Runtime evaluation belongs on named hardware, driver, library, kernel, batch, sequence shape, warm-up, and measurement method. A vendor result cannot be transferred to a different stack by repeating its number.

KV-cache quantization error studies quantized runtime state, which may evolve per request. Static MX blocks and dynamic cache values have different scale lifetimes and failure surfaces. Test both when a deployment uses both; do not let a weight-format audit stand in for a cache audit.

The scalar lab makes no throughput, energy, or kernel-parity claim. Its job is to make logical conversion reviewable. A release note can then connect that receipt to measured file, quality, and runtime evidence without blurring the sources.

Publish the MXFP4 audit receipt

A durable receipt names the 32 inputs, block size, E2M1 table version, E8M0 scale policy, exponent and code, normalized values, nibbles, reconstructions, packed lab serialization, counts, errors, input hash, and schema version. It also repeats the 136-bit and 4.25-bit derivation and the physical-layout disclaimer.

Reject incomplete blocks, non-finite values, magnitudes above 2^120, nonzero block maxima below 2^-125, exponents outside −127 through 127, and policy names outside the fixed set. The lower input bound keeps the max-based result within E8M0's finite range; the separate all-zero policy remains explicit. Byte 255 is the E8M0 NaN code and is never produced by a finite exponent; exponent −127 maps to byte 0 and exponent 127 maps to byte 254. Bounded inputs keep the lab deterministic and make hostile cases ordinary tests instead of browser surprises. No network request, dynamic code, or unsafe HTML sink is necessary.

Revisit this MXFP4 quantization guide on 2027-01-29, or sooner if OCP revises MX, gpt-oss packaging changes, or the cited TorchAO/TensorRT path changes. Refresh prose, figures, lab fixtures, and receipt hashes together.

The MXFP4 quantization decision is simple: accept a storage or quality claim only after the block arithmetic, actual container, model evaluation, and runtime measurement are each supported by their own evidence. A four-bit element is only the beginning of that chain.

Runnable local artifact — The lab audits logical MX encodings and a named scalar scale policy; it does not prescribe checkpoint layout, hardware arithmetic, speed, energy, or model quality.

Plain text1 line
Validate exactly 32 finite inputs, choose reference-max or manual E8M0 scaling, round on the E2M1 lattice, pack a labeled lab serialization, and export a canonical receipt.