Fix GPU divide-by-zero for zero-extent samples in Flip and color conversion - #6506
jantonguirao wants to merge 7 commits into
Conversation
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
No unresolved blocking issues were identified.
Review effort: Lite
Findings: None
What changed in this PR
Fixes GPU divide-by-zero failures for zero-extent samples by adding early returns before CUDA launch configuration.
Changes:
- Guards Flip, color conversion, transform, and matrix launchers against zero work.
- Adds regression tests for zero-extent Flip and color conversion cases.
| File | Description |
|---|---|
dali/operators/image/remap/warp_affine_params.cu |
Skips empty transform launches. |
dali/operators/image/remap/cvcuda/matrix_adjust.cu |
Skips empty matrix launches. |
dali/kernels/imgproc/flip_gpu.cuh |
Skips zero-extent Flip launches. |
dali/kernels/imgproc/flip_gpu_test.cu |
Tests mixed zero-extent samples. |
dali/kernels/imgproc/color_manipulation/color_space_conversion_kernel.cuh |
Skips zero-pixel conversions. |
dali/kernels/imgproc/color_manipulation/color_space_conversion_kernel_test.cu |
Tests zero-pixel conversion handling. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
|
|
!build |
|
CI MESSAGE: [69448342]: BUILD STARTED |
|
CI MESSAGE: [69448342]: BUILD PASSED |
|
!build |
|
CI MESSAGE: [70400019]: BUILD STARTED |
|
CI MESSAGE: [70400019]: BUILD FAILED |
|
!build |
|
CI MESSAGE: [70414726]: BUILD STARTED |
|
CI MESSAGE: [70414726]: BUILD FAILED |
|
!build |
|
CI MESSAGE: [70432157]: BUILD STARTED |
|
CI MESSAGE: [70432157]: BUILD FAILED |
…UDA affine remap InvertTransforms/CopyTransforms and adjustMatrices compute their CUDA launch grid by dividing by a batch-derived count. When the batch is empty, that count is zero, causing a host-side division by zero before any kernel is even launched. Return early from each launcher when there is no work to do. RunColorSpaceConversionKernel and FlipImpl have the same class of bug for zero-extent samples within an otherwise non-empty batch, but that part is already fixed in NVIDIA#6505, so this is scoped to the two remaining call sites. Fixes NVIDIA#6504
Cover the two remaining early-return guards in this PR: - CopyTransformsGPU<ndims, invert> (warp_affine_params.cu) with count == 0 - adjustMatrices (matrix_adjust.cu) with a zero-batch nvcv::Tensor Both used to divide by zero while computing the CUDA launch grid before any kernel was launched.
matrix_adjust_test.cu includes nvcv/Tensor.hpp directly, but dali_operator_test only links dali_operators PUBLIC, and dali_operators links cvcuda PRIVATE, so the include directories never propagate to the test target, breaking the build with BUILD_CVCUDA=ON.
dali_operators is built with -fvisibility=hidden, so the explicit CopyTransformsGPU specializations were GLOBAL HIDDEN symbols, invisible to dali_operator_test which links against dali_operators as a separate binary. warp_affine_params_test.cu calls these directly, causing an "undefined reference" link failure in CI. Mark the declaration DLL_PUBLIC to export it.
The zero-count/zero-batch tests only assert that the launchers don't crash, which a no-op or broken kernel would also satisfy. Add a companion case to each that runs a non-empty batch through the same launcher and checks the output against a hand-computed reference.
5d159c7 to
0d6c136
Compare
|
!build |
|
CI MESSAGE: [70468688]: BUILD STARTED |
Same visibility issue as CopyTransformsGPU: dali_operators is built with -fvisibility=hidden, so adjustMatrices and nvcvop::GetDataType(DALIDataType, int) were hidden symbols, invisible to dali_operator_test. matrix_adjust_test.cu calls both directly (the latter via the GetDataType<T>() template), causing an "undefined reference" link failure in CI. Mark both declarations DLL_PUBLIC to export them. Verified via readelf that both go from GLOBAL HIDDEN to GLOBAL DEFAULT, and that the resulting objects link cleanly against dali_operators/nvcv/cvcuda.
|
CI MESSAGE: [70468688]: BUILD FAILED |
|
!build |
|
CI MESSAGE: [70479418]: BUILD STARTED |
Matches the repo's DALI memory API convention and releases the buffers automatically if a CUDA call or assertion fails partway through.
|
!build |
|
CI MESSAGE: [70488962]: BUILD STARTED |
|
CI MESSAGE: [70479418]: BUILD FAILED |
|
CI MESSAGE: [70488962]: BUILD PASSED |
Category:
Bug fix (non-breaking change which fixes an issue)
Description:
Fixes #6504.
InvertTransforms/CopyTransformsandadjustMatricescompute their CUDA launchgrid by dividing by a batch-derived count. When the batch is empty, that count is zero, causing
a host-side division by zero before any kernel is even launched. The fix returns early from each
of these launchers when there is no work to do.
RunColorSpaceConversionKernelandFlipImplhave the same class of bug for zero-extentsamples within an otherwise non-empty batch, but that part is already fixed in #6505 — this PR
is scoped to the two remaining call sites to avoid duplicating that fix.
Additional information:
Affected modules and functionalities:
dali/operators/image/remap/warp_affine_params.cu:InvertTransforms/CopyTransformsnowreturn early when
count == 0.dali/operators/image/remap/cvcuda/matrix_adjust.cu:adjustMatricesnow returns early whenthe batch size is
0.Key points relevant for the review:
Both fixed sites follow the same shape: a host-side
<<<grid, block>>>launch configurationcomputed by dividing by a count that can be zero. The fix is a minimal early-return guard in
each shared launcher, not a behavioral change for any non-zero case.
Tests:
warp_affine_params.cuwas compile-verified.matrix_adjust.cu(CV-CUDA) could not bebuild-verified in this environment and was reviewed only; it follows the same one-line pattern.
Checklist
Documentation
DALI team only
Requirements
REQ IDs: N/A
JIRA TASK: N/A