Skip to content

Guard against zero scale for all-zero blocks in weight quantization - #2851

Open
TusharND12 wants to merge 2 commits into
apple:mainfrom
TusharND12:fix-quant-zero-scale-all-zero-block
Open

TusharND12 wants to merge 2 commits into
apple:mainfrom
TusharND12:fix-quant-zero-scale-all-zero-block

Conversation

@TusharND12

@TusharND12 TusharND12 commented Sep 8, 2026 •

Copy link
Copy Markdown

Fixes #2853, which has a self-contained repro using only public APIs, the output it produces and the output expected.

Summary

Weight quantization divides by zero when a quantization block is all zeros — for example a pruned output channel, which is what prune_weights produces and joint prune + quantize is a supported workflow.

In coremltools/optimize/_utils.py, quantize_weight_by_dtype does:

scale = (val_max - val_min) / (q_val_max - q_val_min)
quantized_data = np.round(weight / scale)
...
zero_point = (q_val_min * val_max - q_val_max * val_min) / (val_max - val_min)

When a block is all zeros, val_max == val_min, so scale == 0. Both weight / scale and the LINEAR zero_point expression evaluate 0 / 0 and produce NaN, which is then cast to the integer weight dtype. Four RuntimeWarnings are emitted and a scale of exactly 0 is written into constexpr_blockwise_shift_scale.

To be accurate about severity: the dequantized values still read back as 0.0 on x86, because a zero scale multiplies the NaN-derived data away — the right answer, but arrived at by accident rather than computed. What is wrong is the divide-by-zero itself, the degenerate scale = 0 serialized into the model, and model contents that depend on NaN -> int cast behaviour, which NumPy documents as undefined.

The scale can also underflow to zero

The same division is reachable without an all-zero block. The scale is computed in the weight dtype, and Core ML weights are fp16 by default, so a channel whose range is small enough underflows to scale == 0 while still holding nonzero values:

channel magnitude 1e-5 -> scale = 1.788e-07, dequantizes correctly
channel magnitude 1e-6 -> scale = 0.0,       "divide by zero encountered in divide", channel -> all zeros

Internal inconsistency

The palettization path already guards exactly this condition, in coremltools/optimize/coreml/_quantization_passes.py:

per_channel_scale[per_channel_scale == 0] = 1

The linear-quantization path does not. This PR brings it in line.

Fix

Guard scale == 0 directly, mirroring the palettization line above, so it covers both the all-zero block and the underflow case. The zero-range guard is kept as well, so the zero_point division stays safe when val_max == val_min. Both guards fire only on values that are exactly 0, so every well-conditioned block takes the original code path unchanged.

Testing

Two regression tests in TestComputeQuantizationParams:

  • test_compute_qparams_all_zero_block, parametrized over LINEAR / LINEAR_SYMMETRIC and per-channel / per-tensor / blockwise granularities. Asserts the scale is never 0, the quantized data is finite, and the all-zero block round-trips back to 0.
  • test_compute_qparams_scale_underflow, covering the fp16 underflow path.

Both turn any divide-by-zero RuntimeWarning into a failure. Both fail on main and pass with this change.

Verified locally on Linux (conversion and compression only, no on-device prediction):

  • coremltools/test/optimize/test_utils.py — 192 passed.
  • coremltools/test/optimize/coreml/test_passes.py — 4 failed, 4928 passed, identical with and without this patch. The four failures are TestGetActivationStats, which requires prediction.
  • End to end through the public linear_quantize_weights API: the all-zero channel now quantizes with no warnings and dequantizes to 0.0; the underflow channel gets scale = 1.0 instead of 0.0, also with no warnings. Non-degenerate channels are unchanged.

Scope

This targets the shared quantize_weight_by_dtype, used by compute_qparams and linear_quantize_weights.

For the underflow case the guard removes the divide-by-zero and the undefined cast, but the channel still quantizes to zeros, since its values are below what an fp16 scale can represent. Preserving them would mean computing the scale in fp32 — a behavioural change I did not want to make unilaterally, but happy to take on if reviewers prefer it.

`quantize_weight_by_dtype` computes
`scale = (val_max - val_min) / (q_val_max - q_val_min)` and then
`np.round(weight / scale)`. When a quantization block is all zeros (e.g. a
pruned output channel), `val_max == val_min`, so the scale is 0 and both
`weight / scale` and the LINEAR zero_point computation divide by zero, silently
producing NaNs -- the resulting scale/zero_point/quantized weights are garbage.

Guard the range against zero (replace a zero `val_max - val_min` with 1) so an
all-zero block gets a valid nonzero scale, quantizes to 0 and dequantizes back
to 0. This mirrors the palettization path, which already guards the same case
with `per_channel_scale[per_channel_scale == 0] = 1` in
coreml/_quantization_passes.py.

Reachable through the public `linear_quantize_weights` compression pass on any
model with an all-zero (pruned) output channel.

Add a regression test covering per-channel, per-tensor and blockwise
granularities for both LINEAR and LINEAR_SYMMETRIC modes; it fails before this
change (divide-by-zero) and passes after.
@TobyRoseman

Copy link
Copy Markdown
Collaborator

Please create a GitHub issue for the problem you are trying to solve here and link it in this PR. In that issue, please clearly demonstrate the problem with a minimal code example that is self-contained and uses only public APIs. Include the output you get and the output you expected.

The scale is computed in the weight dtype, and Core ML weights are fp16
by default, so a channel whose range is small enough underflows to a
scale of 0 even when the channel is not all zeros. `weight / scale` then
divides by zero and the resulting inf/NaN is cast to the integer weight
dtype.

Guard `scale == 0` directly, which is what the palettization path already
does (`per_channel_scale[per_channel_scale == 0] = 1`), in addition to
the existing zero-range guard that keeps the zero_point division safe.
@TusharND12

Copy link
Copy Markdown
Author

Thanks — filed as #2853. It has a self-contained repro that uses only public APIs (ct.convert → ct.optimize.coreml.linear_quantize_weights → decompress_weights / get_weights_metadata), the output I get, and the output I expected. I also stated plainly in it that the dequantized values happen to come back correct on x86 today, because scale == 0 multiplies the NaN-derived data away — the concrete problems are the four RuntimeWarnings, the degenerate scale = 0 serialized into constexpr_blockwise_shift_scale, and model contents produced by a NaN -> int cast, which NumPy documents as undefined.

While writing the repro I found a second way to reach the same division, which changed the patch. The scale is computed in the weight dtype, and Core ML weights are fp16 by default, so a channel whose range is small enough underflows to scale == 0 even though it is not all zeros:

channel magnitude 1e-5 -> scale = 1.788e-07, dequantizes correctly
channel magnitude 1e-6 -> scale = 0.0,       "divide by zero encountered in divide", channel -> all zeros

So the guard belongs on scale == 0 rather than only on val_max == val_min — which is also exactly what the palettization path already does (per_channel_scale[per_channel_scale == 0] = 1). I pushed 4e70c6b to do that, keeping the zero-range guard for the zero_point division, and added a regression test for the underflow case.

Verification on Linux (conversion and compression only, no on-device prediction):

  • coremltools/test/optimize/test_utils.py: 192 passed.
  • coremltools/test/optimize/coreml/test_passes.py: 4 failed, 4928 passed — identical with and without the patch. The four failures are TestGetActivationStats, which needs prediction.
  • Both new tests fail without the patch and pass with it.
  • End to end through the public API: the all-zero channel now quantizes with no warnings and dequantizes to 0.0, and the underflow channel gets scale = 1.0 instead of 0.0 with no warnings. Non-degenerate channels are byte-identical either way.

One thing I deliberately did not change: for the underflow case the guard removes the divide-by-zero and the undefined cast, but the channel still quantizes to zeros, since its values are below what an fp16 scale can represent. Preserving them would mean computing the scale in fp32, which is a behavioural change I did not want to make unilaterally. Happy to take that on if you would prefer it.

@TusharND12

Copy link
Copy Markdown
Author

Hi! Just following up when you have a chance. I created and linked the requested issue with a minimal public-API reproduction, and updated the patch and regression coverage based on what that investigation uncovered. Please let me know if any additional changes or tests would be helpful. Thanks!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

linear_quantize_weights: divide-by-zero and NaN-to-int cast when a weight block is all zeros

2 participants