Fix AutoEP global L2 clipping for sharded expert ownership - #8476
gss10282023 wants to merge 3 commits into
Conversation
Signed-off-by: shanshan <sgao5239@uni.sydney.edu.au>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e05fa3fa14
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if not isinstance(ep_size, int) or ep_size <= 0 or dp_world_size % ep_size != 0: | ||
| raise RuntimeError("AutoEP expert parameter has invalid EP ownership metadata") | ||
| weight = ep_size / dp_world_size |
There was a problem hiding this comment.
Base expert weighting on the actual reduction group
With mesh-based sequence parallelism, mpu remains None but pg is only the mesh data-parallel group; AutoEP arranges this group as the expert-data-parallel replicas, so every rank in pg owns the same expert shard. Multiplying each shard by ep_size / dp_world_size therefore overstates its squared-norm contribution by ep_size (or raises here when ep_size does not divide the mesh DP size), causing overly aggressive clipping for supported FP32 ZeRO-0 AutoEP+SP runs. The weight must be derived from the shard's replica count within pg, rather than assuming that pg also contains the EP dimension.
Useful? React with 👍 / 👎.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Assisted-by: OpenAI Codex Signed-off-by: shanshan <sgao5239@uni.sydney.edu.au>
Summary
AutoEP FP32 gradient clipping should count each logical parameter once, even when experts are sharded across ranks. Sum ownership-weighted squared gradients before taking the square root, instead of averaging rank-local norms.
For the
mpu=Nonepath, reduce over both data-parallel and sequence-parallel ranks. A mesh DP group can contain only replicas of one expert shard; it does not necessarily cover all experts. Using that smaller group can reject a valid EP size or over-clip gradients. Including all ranks in this path also ensures that the norm contains the other expert shards.Fixes #8475. The change is limited to AutoEP FP32 L2 clipping with
mpu=None; the integration tests use ZeRO-0.Changes
Validation
Tested on an Intel Xeon Cascadelake CPU with Python 3.11.15 and PyTorch 2.13.0+cpu.
CPU/Gloo regression commands
From the repository root in an environment with DeepSpeed's test dependencies:
AI assistance was used for the implementation, tests and this description.