Skip to content

support cuda graph capture offloading module - #2435

Merged
yaox12 merged 35 commits into
NVIDIA:mainfrom
lhb8125:hongbinl/offload_activation_cuda_graph
Mar 2, 2026
Merged

yaox12 merged 35 commits into
NVIDIA:mainfrom
lhb8125:hongbinl/offload_activation_cuda_graph

Conversation

@lhb8125

@lhb8125 lhb8125 commented Dec 1, 2025 •

Copy link
Copy Markdown
Contributor

Description

This PR supports offloading modules captured by partial cuda graph in Megatron-LM.

Fixes # (issue)

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

Please list the changes introduced in this PR:

  • Megatron's offloading relys on cpu_offload_v1 but we need to reuse mark_not_offload() so that the weights won't be offloaded.
  • Do not offload the output of core_attention
  • Refine the allocation strategy of fp8&fp4 tensors because the previous impl allocates tensor by from_blob(), which is not compatitable with record_stream()
  • Add pre_warmup_hook and post_warmup_hook
  • Passing cuda_graph_stream and cuda_graph_event to user_kwargs so that the fwd&bwd replay runs at a side stream, where the cuda_graph_event records on current stream after finishing computing.

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
@lhb8125
lhb8125 marked this pull request as draft December 1, 2025 05:34
@greptile-apps

greptile-apps Bot commented Dec 1, 2025 •

Copy link
Copy Markdown
Contributor

Greptile Summary

Adds CUDA graph capture offloading support for Megatron-LM integration. The PR introduces cuda_graph_stream and cuda_graph_event parameters to enable graph replay on side streams, allowing overlap with main stream work. It also adds pre_warmup_hook and post_warmup_hook callbacks for setup/teardown during graph capture warmup.

Key changes:

  • Modified autograd function signatures to accept cuda_graph_stream and cuda_graph_event parameters
  • Added conditional stream/event synchronization logic in forward and backward passes
  • Hooks are called once before/after warmup loop (not on each iteration)
  • Updated backward return signature to match new forward parameters

Previously raised concerns that remain valid:

  • cuda_graph_event is silently ignored if cuda_graph_stream is not different from current stream, which could lead to synchronization bugs
  • Stream/event extraction from user_kwargs uses pop(), though this is intentional to prevent passing them to underlying functions

The implementation follows the stated goal of supporting offloading with partial CUDA graphs, and the synchronization logic appears sound when parameters are used as intended.

Confidence Score: 3/5

  • Acceptable with some concerns - the stream/event synchronization logic needs validation improvements
  • The implementation adds complex CUDA graph stream/event handling that appears logically sound, but lacks validation for edge cases (e.g., event provided without different stream). Previous review threads identified legitimate concerns about synchronization. The core functionality follows established patterns for CUDA graph capture, and documentation has been added for new parameters.
  • Pay attention to transformer_engine/pytorch/graph.py lines 869-878 (stream/event parameter extraction) and lines 804-812, 830-839 (stream/event synchronization logic)

Important Files Changed

Filename Overview
transformer_engine/pytorch/graph.py Adds CUDA graph capture offloading support with custom stream/event handling, pre/post warmup hooks, and modified forward/backward signatures to support side-stream execution

Last reviewed commit: 5115a1c

@greptile-apps greptile-apps Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 file reviewed, 3 comments

Edit Code Review Agent Settings | Greptile

Comment thread transformer_engine/pytorch/graph.py Outdated
Comment thread transformer_engine/pytorch/graph.py Outdated
Comment thread transformer_engine/pytorch/graph.py Outdated
lhb8125 and others added 17 commits December 8, 2025 06:35
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: root <root@eos0046.eos.clusters.nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
@lhb8125

lhb8125 commented Feb 5, 2026

Copy link
Copy Markdown
Contributor Author

/te-ci pytorch L1

@lhb8125
lhb8125 marked this pull request as ready for review February 5, 2026 03:37

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 files reviewed, no comments

Edit Code Review Agent Settings | Greptile

@lhb8125

lhb8125 commented Feb 5, 2026

Copy link
Copy Markdown
Contributor Author

@buptzyb @zhongbozhu @pggPL Could you review this PR?

"""
return self.wgrad_store is not None and self.wgrad_store.delay_wgrad_compute()

def trigger_backward_dw(self):

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please ignore this method, which will be removed after https://github.com/NVIDIA/TransformerEngine/pull/2614/files merged

Comment thread transformer_engine/pytorch/graph.py Outdated
bwd_dw_graphs[graph_idx].replay()
for module in te_modules:
if hasattr(module, "trigger_backward_dw"):
module.trigger_backward_dw()

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please ignore this code block, which will be removed after https://github.com/NVIDIA/TransformerEngine/pull/2614/files are merged

pool: Optional[Tuple[int, ...]] = None,
retain_graph_in_backward: bool = False,
_reuse_graph_input_output_buffers: bool = False,
pre_warmup_hook: Optional[Callable] = None,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just to confirm: are the hooks used to disable and re-enable offloading for warmup? Could you point me to the implementation in MCore?

@lhb8125 lhb8125 Feb 25, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comment thread transformer_engine/pytorch/graph.py Outdated
Comment thread transformer_engine/pytorch/graph.py Outdated
Comment thread transformer_engine/pytorch/graph.py Outdated
Comment thread transformer_engine/pytorch/graph.py
Comment thread transformer_engine/pytorch/graph.py
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 file reviewed, no comments

Edit Code Review Agent Settings | Greptile

@lhb8125
lhb8125 force-pushed the hongbinl/offload_activation_cuda_graph branch from a181176 to aeb3ecb Compare February 25, 2026 07:08

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 file reviewed, no comments

Edit Code Review Agent Settings | Greptile

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
@lhb8125
lhb8125 force-pushed the hongbinl/offload_activation_cuda_graph branch from aeb3ecb to 60ad9a7 Compare February 25, 2026 07:12
@lhb8125

lhb8125 commented Feb 25, 2026

Copy link
Copy Markdown
Contributor Author

/te-ci pytorch L1

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 file reviewed, no comments

Edit Code Review Agent Settings | Greptile

@lhb8125
lhb8125 requested a review from buptzyb February 25, 2026 07:19
Comment thread transformer_engine/pytorch/graph.py
lhb8125 and others added 2 commits February 25, 2026 18:35
@lhb8125

lhb8125 commented Feb 26, 2026

Copy link
Copy Markdown
Contributor Author

/te-ci pytorch L1

@lhb8125

lhb8125 commented Feb 27, 2026

Copy link
Copy Markdown
Contributor Author

/te-ci pytorch L1

@lhb8125

lhb8125 commented Mar 2, 2026

Copy link
Copy Markdown
Contributor Author

/te-ci pytorch L1

@lhb8125

lhb8125 commented Mar 2, 2026

Copy link
Copy Markdown
Contributor Author

/te-ci pytorch L1

@yaox12 yaox12 left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@yaox12
yaox12 merged commit bba7bf6 into NVIDIA:main Mar 2, 2026
29 of 32 checks passed
phu0ngng pushed a commit to phu0ngng/TransformerEngine that referenced this pull request Mar 2, 2026
* support cuda graph capture offloading module

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* remove reset_hook and init_chunk_handler_hook

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* remove reset_hook and init_chunk_handler_hook

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* minor fix

Signed-off-by: root <root@eos0046.eos.clusters.nvidia.com>

* temp fix overlap-grad-reduce

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* reuse mark_not_offload() and do not offload scale_inv

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* temp fix for mxfp8

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* fix bug for record_stream and from_blob

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* disable offloading core_attn_out and refine cpu overhead of at::empty

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* minor fix

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* return ptr of whole buffer and offload the whole buffer

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Apply suggestions from code revie

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* remove code changes of offloading and quantizer

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* minor fix

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* minor fix

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* minor fix

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* minor fix

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* minor fix

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

* add docstring

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>

---------

Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: root <root@eos0046.eos.clusters.nvidia.com>
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
Co-authored-by: root <root@eos0046.eos.clusters.nvidia.com>
Co-authored-by: root <root@eos0022.eos.clusters.nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants