Skip to content

Compress CUDA binaries - #3465

Merged
ptrendx merged 1 commit into
NVIDIA:mainfrom
fheinecke:fred/compress-cubins-1
Sep 3, 2026
Merged

ptrendx merged 1 commit into
NVIDIA:mainfrom
fheinecke:fred/compress-cubins-1

Conversation

@fheinecke

@fheinecke fheinecke commented Sep 2, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Compresses build CUDA binaries to shrink the TE core library size.

I ran a build with and without this PR. The impact on CUDA 13 TE amd64 builds is significant:

  • The core library size moves from 1058 MB to 361 MB, cutting the size by two thirds
  • The built wheel for main without this change is 556 MB, but implementing this cuts the size to 292 MB (which is much closer to the 2.18 release wheel size despite containing more GPU architectures and features)
  • The built container image is currently 1221 MB, reduced to 719 MB with this changeset

The downsides are:

  • Wheel build time is increased by about 7%
  • The CUDA driver must decompress this data before first use

I'm not expecting decompression time to be meaningful, given that other projects (like pytorch itself) is using this same approach already. It's possible in some systems startup time may actually be improved, if the added decompression time is less than the time savings from reduced disk I/O.

Type of change

  • Documentation change (change only to the documentation, either a fix or a new content)
  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Infra/Build change
  • Code refactoring

Changes

Please list the changes introduced in this PR:

  • Compress build CUDA binaries to reduce TE core library size

Checklist:

  • I have read and followed the contributing guidelines
  • The functionality is complete
  • I have commented my code, particularly in hard-to-understand areas
  • I have made corresponding changes to the documentation
  • My changes generate no new warnings
  • I have added tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Signed-off-by: Fred Heinecke <fheinecke@nvidia.com>
@fheinecke

Copy link
Copy Markdown
Collaborator Author

/te-ci

@greptile-apps

greptile-apps Bot commented Sep 3, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR enables compression for every embedded CUDA fatbinary and selects size-optimized compression for CUDA 13 or newer, reducing the packaged native core library and wheel sizes.

  • Adds -Xfatbin -compress-all to common CUDA compilation.
  • Adds --compress-mode=size when building with CUDA 13 or newer.
  • Retains the legacy/default compression mode for CUDA 12 builds.

Confidence Score: 5/5

The PR appears safe to merge, with no concrete build, compatibility, or runtime defect identified.

The added options use the existing global CUDA flag construction pattern, and size mode is restricted to CUDA 13 or newer while CUDA 12 retains the compatibility-oriented default mode.

Important Files Changed

Filename Overview
transformer_engine/common/CMakeLists.txt Adds version-gated CUDA fatbinary compression flags; no concrete build or runtime regression was established across the repository’s supported configurations.

Reviews (1): Last reviewed commit: "Compress CUDA binaries" | Re-trigger Greptile

@ptrendx
ptrendx merged commit 888da30 into NVIDIA:main Sep 3, 2026
40 of 47 checks passed
fheinecke added a commit that referenced this pull request Sep 3, 2026
Signed-off-by: Fred Heinecke <fheinecke@nvidia.com>
(cherry picked from commit 888da30)
@fheinecke fheinecke added the 2.19 label Sep 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants