Test Copy Engine All-to-all by kwen2501 · Pull Request #170344 · pytorch/pytorch

kwen2501 · 2025-12-12T22:42:43Z

Stack from ghstack (oldest at bottom):

-> Test Copy Engine All-to-all #170344

NCCL 2.28 added Copy Engine (CE) support.

Condition:

Tensors be symmetrically registered (e.g. coming from symm_mem.empty)
NCCL_CTA_POLICY_ZERO be passed to ncclConfig or env var NCCL_CTA_POLICY=2

Confirmed use of CE via profile:

(First kernel is from all_to_all_single on regular tensor, second kernel is from all_to_all_single on tensors that have been window registered)

Caveat:
As of 2.28.9, CE collectives cannot be run on default stream, so we are testing it with async_op=True or with a side stream.

[ghstack-poisoned]

pytorch-bot · 2025-12-12T22:42:46Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/170344

📄 Preview Python docs built from this PR
📄 Preview C++ docs built from this PR
❓ Need help or want to give feedback on the CI? Visit the bot commands wiki

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure

As of commit e212bdf with merge base 8c5e14f ():

NEW FAILURE - The following job has failed:

linux-aarch64 / linux-jammy-aarch64-py3.10 / test (openreg, 1, 1, linux.arm64.m8g.4xlarge) (gh)
'Test'

This comment was automatically generated by Dr. CI and updates every 15 minutes.

ghstack-source-id: 372b50b Pull-Request: #170344

fduwjj · 2025-12-12T23:20:30Z

test/distributed/test_ce_colls.py

+        # if self.rank == 0:
+        #     prof.export_chrome_trace("test_ce_alltoall.json")


It's left here on purpose - when dump of trace is needed from this test :)

No I don't think leave a comment like this makes sense.

[ghstack-poisoned]

ghstack-source-id: 4a8e7fc Pull-Request: #170344

fduwjj · 2025-12-17T01:44:05Z

@pytorchbot merge

pytorchmergebot · 2025-12-17T01:45:53Z

Merge started

Your change will be merged once all checks pass (ETA 0-4 Hours).

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging

Check the merge workflow status
here

pytorchmergebot · 2025-12-17T02:23:01Z

Merge failed

Reason: 1 jobs have failed, first few of them are: trunk / linux-jammy-cuda13.0-py3.10-gcc11 / test (pr_time_benchmarks, 1, 1, linux.g4dn.metal.nvidia.gpu)

Details for Dev Infra team

Raised by workflow job

kwen2501 · 2025-12-17T07:06:11Z

@pytorchbot merge -f "pr_time_benchmark failure is unrelated"

pytorchmergebot · 2025-12-17T07:08:35Z

Merge started

Your change will be merged immediately since you used the force (-f) flag, bypassing any CI checks (ETA: 1-5 minutes). Please use -f as last resort and instead consider -i/--ignore-current to continue the merge ignoring current failures. This will allow currently pending tests to finish and report signal before the merge.

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging

Check the merge workflow status
here

NCCL 2.28 added Copy Engine (CE) support. Condition: - Tensors be symmetrically registered (e.g. coming from symm_mem.empty) - NCCL_CTA_POLICY_ZERO be passed to ncclConfig or env var NCCL_CTA_POLICY=2 Confirmed use of CE via profile: <img width="612" height="167" alt="Screenshot 2025-12-12 at 2 44 23 PM" src="https://github.com/user-attachments/assets/5efb6e9c-40a4-43a0-878f-36733b8b64dd" /> (First kernel is from `all_to_all_single` on regular tensor, second kernel is from `all_to_all_single` on tensors that have been window registered) Caveat: As of 2.28.9, CE collectives cannot be run on default stream, so we are testing it with `async_op=True` or with a side stream. Pull Request resolved: pytorch#170344 Approved by: https://github.com/fduwjj

Update

9e8f840

[ghstack-poisoned]

pytorch-bot bot added the topic: not user facing topic category label Dec 12, 2025

kwen2501 added a commit that referenced this pull request Dec 12, 2025

Test Copy Engine All-to-all

57763f5

ghstack-source-id: 372b50b Pull-Request: #170344

kwen2501 added release notes: distributed (symm_mem) release note label for symmetric memory module: symm_mem Issues and PRs of Symmetric Memory and removed topic: not user facing topic category labels Dec 12, 2025

pytorchbot added the open source label Dec 12, 2025

kwen2501 requested review from fduwjj and ngimel December 12, 2025 22:46

pytorch-bot bot added the topic: not user facing topic category label Dec 12, 2025

fduwjj reviewed Dec 12, 2025

View reviewed changes

Update

e212bdf

[ghstack-poisoned]

kwen2501 added a commit that referenced this pull request Dec 16, 2025

Test Copy Engine All-to-all

495c5cb

ghstack-source-id: 4a8e7fc Pull-Request: #170344

kwen2501 requested a review from fduwjj December 17, 2025 00:48

fduwjj approved these changes Dec 17, 2025

View reviewed changes

pytorch-bot bot added the ciflow/trunk Trigger trunk jobs on your pull request label Dec 17, 2025

pytorchmergebot added the merging label Dec 17, 2025

pytorchmergebot removed the merging label Dec 17, 2025

pytorchmergebot added the merging label Dec 17, 2025

pytorchmergebot added the Merged label Dec 17, 2025

pytorchmergebot closed this in 4b20d75 Dec 17, 2025

pytorchmergebot removed the merging label Dec 17, 2025

github-actions bot deleted the gh/kwen2501/294/head branch January 17, 2026 02:19

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Test Copy Engine All-to-all#170344

Test Copy Engine All-to-all#170344
kwen2501 wants to merge 2 commits intogh/kwen2501/294/basefrom
gh/kwen2501/294/head

kwen2501 commented Dec 12, 2025 •

edited

Loading

Uh oh!

pytorch-bot bot commented Dec 12, 2025 •

edited

Loading

Uh oh!

fduwjj Dec 12, 2025

Uh oh!

kwen2501 Dec 13, 2025

Uh oh!

fduwjj Dec 15, 2025

Uh oh!

kwen2501 Dec 16, 2025

Uh oh!

fduwjj commented Dec 17, 2025

Uh oh!

pytorchmergebot commented Dec 17, 2025

Uh oh!

pytorchmergebot commented Dec 17, 2025

Uh oh!

kwen2501 commented Dec 17, 2025

Uh oh!

pytorchmergebot commented Dec 17, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

		# if self.rank == 0:
		# prof.export_chrome_trace("test_ce_alltoall.json")

Conversation

kwen2501 commented Dec 12, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

pytorch-bot bot commented Dec 12, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/170344

❌ 1 New Failure

Uh oh!

fduwjj Dec 12, 2025

Choose a reason for hiding this comment

Uh oh!

kwen2501 Dec 13, 2025

Choose a reason for hiding this comment

Uh oh!

fduwjj Dec 15, 2025

Choose a reason for hiding this comment

Uh oh!

kwen2501 Dec 16, 2025

Choose a reason for hiding this comment

Uh oh!

fduwjj commented Dec 17, 2025

Uh oh!

pytorchmergebot commented Dec 17, 2025

Merge started

Uh oh!

pytorchmergebot commented Dec 17, 2025

Merge failed

Uh oh!

kwen2501 commented Dec 17, 2025

Uh oh!

pytorchmergebot commented Dec 17, 2025

Merge started

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

4 participants

kwen2501 commented Dec 12, 2025 •

edited

Loading

pytorch-bot bot commented Dec 12, 2025 •

edited

Loading