Skip to content

Fix AutoEP global L2 clipping for sharded expert ownership - #8476

Open
gss10282023 wants to merge 3 commits into
deepspeedai:masterfrom
gss10282023:fix-autoep-ownership-aware-clipping
Open

gss10282023 wants to merge 3 commits into
deepspeedai:masterfrom
gss10282023:fix-autoep-ownership-aware-clipping

Conversation

@gss10282023

@gss10282023 gss10282023 commented Sep 10, 2026 •

Copy link
Copy Markdown

Summary

AutoEP FP32 gradient clipping should count each logical parameter once, even when experts are sharded across ranks. Sum ownership-weighted squared gradients before taking the square root, instead of averaging rank-local norms.

For the mpu=None path, reduce over both data-parallel and sequence-parallel ranks. A mesh DP group can contain only replicas of one expert shard; it does not necessarily cover all experts. Using that smaller group can reject a valid EP size or over-clip gradients. Including all ranks in this path also ensures that the norm contains the other expert shards.

Fixes #8475. The change is limited to AutoEP FP32 L2 clipping with mpu=None; the integration tests use ZeRO-0.

Changes

  • Record the EP size on expert parameters and use the reduction group's size to account for replicas.
  • Detect AutoEP ownership before filtering missing gradients, keeping every rank on the same collective path when an expert is unused.
  • Extend the native clipping/update regression to ordinary and mesh-SP layouts on two and four ranks. Both layouts use the same independent full-gradient reference.

Validation

Tested on an Intel Xeon Cascadelake CPU with Python 3.11.15 and PyTorch 2.13.0+cpu.

  • 32 native engine clipping/update cases passed on CPU/Gloo: 8 ordinary and 8 mesh-SP cases at each of two and four ranks. They cover EP sizes 1 and 2, mixed dense/expert gradients, and an unused expert.
  • Before the mesh-SP correction, the two-rank case raised an invalid-ownership-metadata error. The four-rank case clipped nonzero expert gradients to about 70.7% of the reference value. The ordinary-layout controls passed.
  • The parameter-metadata test and existing p-norm clipping tests pass: 4 passed, 2 skipped. The two distributed test methods are skipped by the standard CPU harness because it sees one accelerator; their bodies were run directly at both world sizes above.
  • All pre-commit checks applicable to the changed files passed. CUDA/NCCL tests have not been run locally for this update.
CPU/Gloo regression commands

From the repository root in an environment with DeepSpeed's test dependencies:

cat > /tmp/autoep_clip_check.py <<'PYTEST'
import deepspeed
import deepspeed.comm as dist
from unit.v1.moe.test_autoep_grad_parity import TestAutoEPFP32Clipping

deepspeed.init_distributed(dist_backend="gloo", auto_mpi_discovery=False)
try:
    test = TestAutoEPFP32Clipping()
    test.test_clipping_counts_each_unique_parameter_once()
    test.test_clipping_includes_expert_shards_across_mesh_sp_ranks()
finally:
    dist.destroy_process_group()
PYTEST

DS_ACCELERATOR=cpu PYTHONPATH=".:tests" OMP_NUM_THREADS=1 \
python -m torch.distributed.run --standalone --nproc-per-node=2 /tmp/autoep_clip_check.py
DS_ACCELERATOR=cpu PYTHONPATH=".:tests" OMP_NUM_THREADS=1 \
python -m torch.distributed.run --standalone --nproc-per-node=4 /tmp/autoep_clip_check.py

AI assistance was used for the implementation, tests and this description.

Signed-off-by: shanshan <sgao5239@uni.sydney.edu.au>
@gss10282023
gss10282023 marked this pull request as ready for review September 10, 2026 16:54

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e05fa3fa14

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread deepspeed/runtime/utils.py Outdated
Comment on lines +406 to +408
if not isinstance(ep_size, int) or ep_size <= 0 or dp_world_size % ep_size != 0:
raise RuntimeError("AutoEP expert parameter has invalid EP ownership metadata")
weight = ep_size / dp_world_size

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Base expert weighting on the actual reduction group

With mesh-based sequence parallelism, mpu remains None but pg is only the mesh data-parallel group; AutoEP arranges this group as the expert-data-parallel replicas, so every rank in pg owns the same expert shard. Multiplying each shard by ep_size / dp_world_size therefore overstates its squared-norm contribution by ep_size (or raises here when ep_size does not divide the mesh DP size), causing overly aggressive clipping for supported FP32 ZeRO-0 AutoEP+SP runs. The weight must be derived from the shard's replica count within pg, rather than assuming that pg also contains the EP dimension.

Useful? React with 👍 / 👎.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-10T16:59:14.542081Z e05fa3f Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

Assisted-by: OpenAI Codex
Signed-off-by: shanshan <sgao5239@uni.sydney.edu.au>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] AutoEP FP32 gradient clipping averages norms across ranks with different expert ownership

1 participant