Support building with headers from nvidia wheels - #2623
Conversation
There are two changes:
1. `import nvidia` returns a namespace package with `__file__` equal to `None`
2. Add the way to force headers from nvidia wheels. Without that envvar, it's practically impossible with CUDA installed system-wide.
I successfully built the package with torch using the following `uv` configuration:
```
[tool.uv.extra-build-dependencies]
"transformer-engine-torch" = [
"ninja",
"nvidia-cuda-crt==13.0.88",
"nvidia-cuda-cccl==13.0.85",
{ requirement = "torch", match-runtime = true },
{ requirement = "pytorch-triton", match-runtime = true },
{ requirement = "nvidia-cusolver", match-runtime = true },
{ requirement = "nvidia-curand", match-runtime = true },
{ requirement = "nvidia-cublas", match-runtime = true },
{ requirement = "nvidia-cusparse", match-runtime = true },
{ requirement = "nvidia-cudnn-cu13", match-runtime = true },
{ requirement = "nvidia-nvtx", match-runtime = true },
{ requirement = "nvidia-cuda-nvrtc", match-runtime = true },
{ requirement = "nvidia-cuda-runtime", match-runtime = true },
]
```
Signed-off-by: Vadim Markovtsev <vadim@poolside.ai>
Greptile OverviewGreptile SummaryEnhanced CUDA header discovery to support building with NVIDIA wheels in environments where CUDA toolkit is installed system-wide. The PR introduces two key improvements:
The changes enable building with specific CUDA component versions from wheels (like Confidence Score: 4/5
Important Files Changed
Sequence DiagramsequenceDiagram
participant Build as Build System
participant Utils as get_cuda_include_dirs()
participant Env as Environment
participant FS as File System
participant Nvidia as nvidia package
Build->>Utils: Request CUDA include dirs
Utils->>Env: Check NVTE_BUILD_USE_NVIDIA_WHEELS
alt force_wheels=True OR toolkit not found
Utils->>Nvidia: import nvidia
alt nvidia.__file__ exists
Nvidia-->>Utils: Return __file__ path
Utils->>FS: cuda_root = Path(__file__).parent
else nvidia.__file__ is None (namespace pkg)
Nvidia-->>Utils: Return None
Utils->>FS: cuda_root = Path(__path__[0])
end
Utils->>FS: Find subdirs with include/
FS-->>Utils: Return list of include dirs
Utils-->>Build: Return wheel-based includes
else force_wheels=False AND toolkit found
Utils->>FS: cuda_toolkit_include_path()
FS-->>Utils: Return toolkit include path
Utils-->>Build: Return toolkit includes
end
|
|
/te-ci L0 |
|
It looks like the red jobs are unrelated timeouts. |
Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com> Signed-off-by: Kirthi Shankar Sivamani <ksivamani@nvidia.com>
|
/te-ci L0 |
| def get_cuda_include_dirs() -> Tuple[str, str]: | ||
| """Returns the CUDA header directory.""" | ||
|
|
||
| force_wheels = bool(int(os.getenv("NVTE_BUILD_USE_NVIDIA_WHEELS", "0"))) |
There was a problem hiding this comment.
Consider documenting the new NVTE_BUILD_USE_NVIDIA_WHEELS environment variable in docs/envvars.rst under the "CUDA Configuration" section for consistency with other build-time environment variables.
Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
There are two changes:
import nvidiareturns a namespace package with__file__equal toNoneI successfully built the package with torch using the following
uvconfiguration:Description
Please include a brief summary of the changes, relevant motivation and context.
Fixes # (issue)
Type of change
Checklist: