support cuda graph capture offloading module - #2435
Conversation
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
for more information, see https://pre-commit.ci
Greptile SummaryAdds CUDA graph capture offloading support for Megatron-LM integration. The PR introduces Key changes:
Previously raised concerns that remain valid:
The implementation follows the stated goal of supporting offloading with partial CUDA graphs, and the synchronization logic appears sound when parameters are used as intended. Confidence Score: 3/5
Important Files Changed
Last reviewed commit: 5115a1c |
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
…ub.com/lhb8125/TransformerEngine into hongbinl/offload_activation_cuda_graph
…ub.com/lhb8125/TransformerEngine into hongbinl/offload_activation_cuda_graph
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
…ub.com/lhb8125/TransformerEngine into hongbinl/offload_activation_cuda_graph
for more information, see https://pre-commit.ci
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
…ub.com/lhb8125/TransformerEngine into hongbinl/offload_activation_cuda_graph
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
for more information, see https://pre-commit.ci
|
/te-ci pytorch L1 |
|
@buptzyb @zhongbozhu @pggPL Could you review this PR? |
| """ | ||
| return self.wgrad_store is not None and self.wgrad_store.delay_wgrad_compute() | ||
|
|
||
| def trigger_backward_dw(self): |
There was a problem hiding this comment.
Please ignore this method, which will be removed after https://github.com/NVIDIA/TransformerEngine/pull/2614/files merged
| bwd_dw_graphs[graph_idx].replay() | ||
| for module in te_modules: | ||
| if hasattr(module, "trigger_backward_dw"): | ||
| module.trigger_backward_dw() |
There was a problem hiding this comment.
Please ignore this code block, which will be removed after https://github.com/NVIDIA/TransformerEngine/pull/2614/files are merged
| pool: Optional[Tuple[int, ...]] = None, | ||
| retain_graph_in_backward: bool = False, | ||
| _reuse_graph_input_output_buffers: bool = False, | ||
| pre_warmup_hook: Optional[Callable] = None, |
There was a problem hiding this comment.
Just to confirm: are the hooks used to disable and re-enable offloading for warmup? Could you point me to the implementation in MCore?
There was a problem hiding this comment.
There was a problem hiding this comment.
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
a181176 to
aeb3ecb
Compare
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
aeb3ecb to
60ad9a7
Compare
|
/te-ci pytorch L1 |
Signed-off-by: Hongbin Liu <hongbinl@nvidia.com>
|
/te-ci pytorch L1 |
|
/te-ci pytorch L1 |
|
/te-ci pytorch L1 |
|
/te-ci pytorch L1 |
* support cuda graph capture offloading module Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * remove reset_hook and init_chunk_handler_hook Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * remove reset_hook and init_chunk_handler_hook Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * minor fix Signed-off-by: root <root@eos0046.eos.clusters.nvidia.com> * temp fix overlap-grad-reduce Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * reuse mark_not_offload() and do not offload scale_inv Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * temp fix for mxfp8 Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * fix bug for record_stream and from_blob Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * disable offloading core_attn_out and refine cpu overhead of at::empty Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * minor fix Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * return ptr of whole buffer and offload the whole buffer Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Apply suggestions from code revie Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * remove code changes of offloading and quantizer Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * minor fix Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * minor fix Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * minor fix Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * minor fix Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * minor fix Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> * add docstring Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> --------- Signed-off-by: Hongbin Liu <hongbinl@nvidia.com> Signed-off-by: root <root@eos0046.eos.clusters.nvidia.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: root <root@eos0046.eos.clusters.nvidia.com> Co-authored-by: root <root@eos0022.eos.clusters.nvidia.com>
Description
This PR supports offloading modules captured by partial cuda graph in Megatron-LM.
Fixes # (issue)
Type of change
Changes
Please list the changes introduced in this PR:
record_stream()pre_warmup_hookandpost_warmup_hookcuda_graph_streamandcuda_graph_eventtouser_kwargsso that the fwd&bwd replay runs at a side stream, where thecuda_graph_eventrecords on current stream after finishing computing.Checklist: