Repository navigation
Guard against zero scale for all-zero blocks in weight quantization - #2851
TusharND12 wants to merge 2 commits into
Conversation
`quantize_weight_by_dtype` computes `scale = (val_max - val_min) / (q_val_max - q_val_min)` and then `np.round(weight / scale)`. When a quantization block is all zeros (e.g. a pruned output channel), `val_max == val_min`, so the scale is 0 and both `weight / scale` and the LINEAR zero_point computation divide by zero, silently producing NaNs -- the resulting scale/zero_point/quantized weights are garbage. Guard the range against zero (replace a zero `val_max - val_min` with 1) so an all-zero block gets a valid nonzero scale, quantizes to 0 and dequantizes back to 0. This mirrors the palettization path, which already guards the same case with `per_channel_scale[per_channel_scale == 0] = 1` in coreml/_quantization_passes.py. Reachable through the public `linear_quantize_weights` compression pass on any model with an all-zero (pruned) output channel. Add a regression test covering per-channel, per-tensor and blockwise granularities for both LINEAR and LINEAR_SYMMETRIC modes; it fails before this change (divide-by-zero) and passes after.
|
Please create a GitHub issue for the problem you are trying to solve here and link it in this PR. In that issue, please clearly demonstrate the problem with a minimal code example that is self-contained and uses only public APIs. Include the output you get and the output you expected. |
The scale is computed in the weight dtype, and Core ML weights are fp16 by default, so a channel whose range is small enough underflows to a scale of 0 even when the channel is not all zeros. `weight / scale` then divides by zero and the resulting inf/NaN is cast to the integer weight dtype. Guard `scale == 0` directly, which is what the palettization path already does (`per_channel_scale[per_channel_scale == 0] = 1`), in addition to the existing zero-range guard that keeps the zero_point division safe.
|
Thanks — filed as #2853. It has a self-contained repro that uses only public APIs ( While writing the repro I found a second way to reach the same division, which changed the patch. The scale is computed in the weight dtype, and Core ML weights are fp16 by default, so a channel whose range is small enough underflows to So the guard belongs on Verification on Linux (conversion and compression only, no on-device prediction):
One thing I deliberately did not change: for the underflow case the guard removes the divide-by-zero and the undefined cast, but the channel still quantizes to zeros, since its values are below what an fp16 scale can represent. Preserving them would mean computing the scale in fp32, which is a behavioural change I did not want to make unilaterally. Happy to take that on if you would prefer it. |
|
Hi! Just following up when you have a chance. I created and linked the requested issue with a minimal public-API reproduction, and updated the patch and regression coverage based on what that investigation uncovered. Please let me know if any additional changes or tests would be helpful. Thanks! |
Fixes #2853, which has a self-contained repro using only public APIs, the output it produces and the output expected.
Summary
Weight quantization divides by zero when a quantization block is all zeros — for example a pruned output channel, which is what
prune_weightsproduces and joint prune + quantize is a supported workflow.In
coremltools/optimize/_utils.py,quantize_weight_by_dtypedoes:When a block is all zeros,
val_max == val_min, soscale == 0. Bothweight / scaleand theLINEARzero_pointexpression evaluate0 / 0and produceNaN, which is then cast to the integer weight dtype. FourRuntimeWarnings are emitted and ascaleof exactly0is written intoconstexpr_blockwise_shift_scale.To be accurate about severity: the dequantized values still read back as
0.0on x86, because a zero scale multiplies theNaN-derived data away — the right answer, but arrived at by accident rather than computed. What is wrong is the divide-by-zero itself, the degeneratescale = 0serialized into the model, and model contents that depend onNaN -> intcast behaviour, which NumPy documents as undefined.The scale can also underflow to zero
The same division is reachable without an all-zero block. The scale is computed in the weight dtype, and Core ML weights are fp16 by default, so a channel whose range is small enough underflows to
scale == 0while still holding nonzero values:Internal inconsistency
The palettization path already guards exactly this condition, in
coremltools/optimize/coreml/_quantization_passes.py:The linear-quantization path does not. This PR brings it in line.
Fix
Guard
scale == 0directly, mirroring the palettization line above, so it covers both the all-zero block and the underflow case. The zero-range guard is kept as well, so thezero_pointdivision stays safe whenval_max == val_min. Both guards fire only on values that are exactly0, so every well-conditioned block takes the original code path unchanged.Testing
Two regression tests in
TestComputeQuantizationParams:test_compute_qparams_all_zero_block, parametrized overLINEAR/LINEAR_SYMMETRICand per-channel / per-tensor / blockwise granularities. Asserts the scale is never0, the quantized data is finite, and the all-zero block round-trips back to0.test_compute_qparams_scale_underflow, covering the fp16 underflow path.Both turn any divide-by-zero
RuntimeWarninginto a failure. Both fail onmainand pass with this change.Verified locally on Linux (conversion and compression only, no on-device prediction):
coremltools/test/optimize/test_utils.py— 192 passed.coremltools/test/optimize/coreml/test_passes.py— 4 failed, 4928 passed, identical with and without this patch. The four failures areTestGetActivationStats, which requires prediction.linear_quantize_weightsAPI: the all-zero channel now quantizes with no warnings and dequantizes to0.0; the underflow channel getsscale = 1.0instead of0.0, also with no warnings. Non-degenerate channels are unchanged.Scope
This targets the shared
quantize_weight_by_dtype, used bycompute_qparamsandlinear_quantize_weights.For the underflow case the guard removes the divide-by-zero and the undefined cast, but the channel still quantizes to zeros, since its values are below what an fp16 scale can represent. Preserving them would mean computing the scale in fp32 — a behavioural change I did not want to make unilaterally, but happy to take on if reviewers prefer it.