Skip to content

fix(kernels): build on rolling-release toolchains (CUDA_PATH veto, POSIX macro clash) - #313

Open
ManuelFCastillo wants to merge 1 commit into
Lightricks:mainfrom
ManuelFCastillo:fix/kernels-rolling-release-build
Open

ManuelFCastillo wants to merge 1 commit into
Lightricks:mainfrom
ManuelFCastillo:fix/kernels-rolling-release-build

Conversation

@ManuelFCastillo

Copy link
Copy Markdown

Fixes #292. Two independent build failures of ltx-kernels on rolling-release toolchains (Arch / CachyOS: GCC 16, glibc 2.42+), each fixed at the cause rather than by loosening -Werror.

1. CUDA_PATH alone no longer vetoes the pip toolkit (setup.py)

_prefer_pip_cuda_home() bailed out if either CUDA_HOME or CUDA_PATH was set. Arch and CachyOS export CUDA_PATH=/opt/cuda globally, so the build silently used the system toolkit against the pinned nvidia-cuda-cccl==13.2.* headers and CCCL aborted with CUDA compiler and CUDA toolkit headers are incompatible.

Now: an explicit CUDA_HOME is still honoured as-is. A bare CUDA_PATH is overridden when the pip nvidia-cuda-nvcc toolkit is installed (both CUDA_HOME and CUDA_PATH are pointed at it), and left alone as the fallback when it is not.

2. torch/python.h leads the include list in all2all.cpp

all2all.cpp included <pybind11/functional.h> after the ATen headers. ATen pulls in glibc <features.h>, which under _GNU_SOURCE now defines _POSIX_C_SOURCE 202405L; pyconfig.h then redefines it to 200809L unconditionally. GCC reports that as a -W-group-less "redefined" warning, so -Werror makes it fatal and -Wno-error=<name> cannot exempt it.

torch/csrc/python_headers.h wraps <Python.h> in push_macro / undef / pop_macro, so it is safe in any position. pybind11's wrap_include_python_h.h does not (and documents that it must precede any standard header). Moving <torch/python.h> to the top makes <Python.h> land first, exactly as CPython's C-API docs require. -Werror is kept.

Verification

I don't have a CachyOS host, so I isolated the mechanism in a minimal include-order matrix compiled with the extension's flags (-Wall -Wextra -Werror) on Fedora 44 (GCC 16.2.1, glibc 2.43, Python 3.14.7, pybind11 3.1.0 from pip via -I, as torch.utils.cpp_extension passes it):

Order Result
<cstdint> → <Python.h> FAIL: '_POSIX_C_SOURCE' redefined [-Werror], '_XOPEN_SOURCE' redefined [-Werror]
<Python.h> → <cstdint> PASS
<cstdint> → <pybind11/pybind11.h> (shipped all2all.cpp order) FAIL: same two diagnostics as the report
<cstdint> → torch/csrc/python_headers.h → pybind11 PASS
torch/csrc/python_headers.h → <cstdint> → pybind11 (this PR's order) PASS

One trap worth noting for anyone reproducing: a distro-packaged pybind11 under /usr/include is a system header dir, and GCC suppresses diagnostics reached through system headers, so the failing case looks fine. The pip wheel reached via -I shows the real behaviour.

The setup.py change was checked with ruff check / ruff format --check and a mocked test of the five CUDA_HOME / CUDA_PATH / pip-present combinations.

Related: #264 relaxes -Werror for a different third-party warning on the same extension; this PR is independent of it.

…SIX macro clash)

Fixes two independent build failures of ltx-kernels on Arch/CachyOS-class
toolchains (GCC 16, glibc 2.42+), reported in Lightricks#292.

1. setup.py only preferred the pip nvidia-cuda-nvcc toolkit when neither
   CUDA_HOME nor CUDA_PATH was set. Arch and CachyOS export CUDA_PATH=/opt/cuda
   globally, so the build silently used the system toolkit against the pinned
   13.2 cccl headers and CCCL aborted with "CUDA compiler and CUDA toolkit
   headers are incompatible". An explicit CUDA_HOME is still honoured; a bare
   CUDA_PATH no longer vetoes the pip toolkit, and is only used as a fallback
   when the pip toolkit is absent.

2. all2all.cpp included <pybind11/functional.h> after the ATen headers. ATen
   pulls in glibc <features.h>, which under _GNU_SOURCE now sets
   _POSIX_C_SOURCE to 202405L; pyconfig.h then redefines it to 200809L
   unconditionally. GCC reports that as an unnamed "redefined" warning, which
   the extension's -Werror turns fatal and -Wno-error=<group> cannot exempt.
   torch/python.h wraps <Python.h> in push_macro/undef/pop_macro, pybind11's
   wrapper does not, so torch/python.h now leads the include list. -Werror is
   kept.

Verified with a minimal include-order matrix on Fedora 44 (GCC 16.2, glibc
2.43, Python 3.14, pybind11 3.1 via -I): the shipped order reproduces the exact
diagnostics from the report, the new order compiles clean under -Werror.

Fixes Lightricks#292
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

ltx-kernels build fails on rolling-release toolchains (CUDA toolkit mismatch + -Werror on glibc/Python POSIX macro clash)

1 participant