fix(kernels): build on rolling-release toolchains (CUDA_PATH veto, POSIX macro clash) - #313
Open
ManuelFCastillo wants to merge 1 commit into
Open
ManuelFCastillo wants to merge 1 commit into
ManuelFCastillo wants to merge 1 commit into
Conversation
…SIX macro clash) Fixes two independent build failures of ltx-kernels on Arch/CachyOS-class toolchains (GCC 16, glibc 2.42+), reported in Lightricks#292. 1. setup.py only preferred the pip nvidia-cuda-nvcc toolkit when neither CUDA_HOME nor CUDA_PATH was set. Arch and CachyOS export CUDA_PATH=/opt/cuda globally, so the build silently used the system toolkit against the pinned 13.2 cccl headers and CCCL aborted with "CUDA compiler and CUDA toolkit headers are incompatible". An explicit CUDA_HOME is still honoured; a bare CUDA_PATH no longer vetoes the pip toolkit, and is only used as a fallback when the pip toolkit is absent. 2. all2all.cpp included <pybind11/functional.h> after the ATen headers. ATen pulls in glibc <features.h>, which under _GNU_SOURCE now sets _POSIX_C_SOURCE to 202405L; pyconfig.h then redefines it to 200809L unconditionally. GCC reports that as an unnamed "redefined" warning, which the extension's -Werror turns fatal and -Wno-error=<group> cannot exempt. torch/python.h wraps <Python.h> in push_macro/undef/pop_macro, pybind11's wrapper does not, so torch/python.h now leads the include list. -Werror is kept. Verified with a minimal include-order matrix on Fedora 44 (GCC 16.2, glibc 2.43, Python 3.14, pybind11 3.1 via -I): the shipped order reproduces the exact diagnostics from the report, the new order compiles clean under -Werror. Fixes Lightricks#292
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #292. Two independent build failures of
ltx-kernelson rolling-release toolchains (Arch / CachyOS: GCC 16, glibc 2.42+), each fixed at the cause rather than by loosening-Werror.1.
CUDA_PATHalone no longer vetoes the pip toolkit (setup.py)_prefer_pip_cuda_home()bailed out if eitherCUDA_HOMEorCUDA_PATHwas set. Arch and CachyOS exportCUDA_PATH=/opt/cudaglobally, so the build silently used the system toolkit against the pinnednvidia-cuda-cccl==13.2.*headers and CCCL aborted withCUDA compiler and CUDA toolkit headers are incompatible.Now: an explicit
CUDA_HOMEis still honoured as-is. A bareCUDA_PATHis overridden when the pipnvidia-cuda-nvcctoolkit is installed (bothCUDA_HOMEandCUDA_PATHare pointed at it), and left alone as the fallback when it is not.2.
torch/python.hleads the include list inall2all.cppall2all.cppincluded<pybind11/functional.h>after the ATen headers. ATen pulls in glibc<features.h>, which under_GNU_SOURCEnow defines_POSIX_C_SOURCE 202405L;pyconfig.hthen redefines it to200809Lunconditionally. GCC reports that as a-W-group-less "redefined" warning, so-Werrormakes it fatal and-Wno-error=<name>cannot exempt it.torch/csrc/python_headers.hwraps<Python.h>inpush_macro/undef/pop_macro, so it is safe in any position. pybind11'swrap_include_python_h.hdoes not (and documents that it must precede any standard header). Moving<torch/python.h>to the top makes<Python.h>land first, exactly as CPython's C-API docs require.-Werroris kept.Verification
I don't have a CachyOS host, so I isolated the mechanism in a minimal include-order matrix compiled with the extension's flags (
-Wall -Wextra -Werror) on Fedora 44 (GCC 16.2.1, glibc 2.43, Python 3.14.7, pybind11 3.1.0 from pip via-I, astorch.utils.cpp_extensionpasses it):<cstdint>→<Python.h>'_POSIX_C_SOURCE' redefined [-Werror],'_XOPEN_SOURCE' redefined [-Werror]<Python.h>→<cstdint><cstdint>→<pybind11/pybind11.h>(shippedall2all.cpporder)<cstdint>→torch/csrc/python_headers.h→ pybind11torch/csrc/python_headers.h→<cstdint>→ pybind11 (this PR's order)One trap worth noting for anyone reproducing: a distro-packaged pybind11 under
/usr/includeis a system header dir, and GCC suppresses diagnostics reached through system headers, so the failing case looks fine. The pip wheel reached via-Ishows the real behaviour.The
setup.pychange was checked withruff check/ruff format --checkand a mocked test of the fiveCUDA_HOME/CUDA_PATH/ pip-present combinations.Related: #264 relaxes
-Werrorfor a different third-party warning on the same extension; this PR is independent of it.