You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Reduce the Philox RNG state kept live across optixTrace() calls in the
photon-propagation loop.
Before tracing, the kernel records the current random-stream position as a
32-bit value:
(rng.ctr.x << 2) | (rng.STATE & 3u)
After tracing, it reconstructs the 64-byte Philox state from the seed, event
index, and photon index, then uses skipahead() to resume at the saved draw
position.
Philox is Simphony's sole device RNG, so this behavior is unconditional.
Motivation
Values that remain live across optixTrace() contribute to the OptiX
continuation-state footprint. Carrying the complete 64-byte Philox state can
therefore increase stack spills and L2-memory traffic.
Keeping only a 4-byte stream position across the trace trades inexpensive
Philox reconstruction for a smaller live set in the propagation kernel.
Correctness
The reconstructed state resumes at the same scalar draw position because:
photon index selects the Philox subsequence;
event index supplies the configured event offset;
ctr.x and STATE identify the current position within the generated
four-value block;
optixTrace() does not consume random values; and
propagation currently uses scalar uniform draws and does not depend on
cached normal-distribution state.
The 32-bit position delta assumes fewer than 2^32 random draws per photon,
which is far above the configured propagation limits.
Fixed-seed validation produced byte-identical hit output between the base and
optimized kernels.
Performance
On the reported DUNE-scale benchmark—125 million photons on an RTX 4090—the
kernel wall time improved by approximately 4%.
Additional pfRICH measurements on an RTX 4090 showed approximately 1.4–1.8%
lower propagation-kernel time, indicating that the exact improvement depends
on geometry and workload.
The reason will be displayed to describe this comment to others. Learn more.
🔵 Needs a closer look
Philox block-boundary handling can restore four draws early and must be corrected.
Pull request overview
This PR reduces OptiX continuation-state pressure by rebuilding Philox RNG state after photon-propagation traces.
Changes:
Saves a compact RNG position before tracing.
Reinitializes and skips ahead after tracing.
Applies the optimization to photon propagation.
File summaries
File
Description
CSGOptiX/CSGOptiX7.cu
Rebuilds Philox RNG state around propagation traces.
Review details
Suppressed comments (1)
CSGOptiX/CSGOptiX7.cu:457
Philox uses a lazy four-value buffer: after the fourth scalar curand_uniform, STATE can be 4 while ctr.x still names the consumed block; the next RNG call normalizes it and advances the counter. Masking STATE to 3 therefore records the beginning of that block, so a bounce ending on this boundary is restored four draws early and repeats random values. Preserve the boundary value (or normalize it while incrementing ctr.x) in both position calculations.
I reran the DUNE benchmarks after confirming that both RTX 4090s were idle. The uncontended results do not reproduce the previously reported speedup; instead, the PR is consistently about 1% slower.
Setup
Base: origin/main at 681f09409
Head: bc3cf5901
Release builds
Fixed seed: 42
DUNE geometry: dune_mock_wls_detector_box.gdml
Eight runs per revision:
Five base → head pairs
Three reverse-order head → base pairs to control for warm-up and thermal ordering
Results
Test
Base median
Head median
Head change
simg4ox, 1M photons
40.257 ms
40.719 ms
1.148% slower
GPURaytrace, summed 122 GPU slices
23.8660 s
24.0880 s
0.930% slower
GPURaytrace, complete simulate()
24.2023 s
24.4241 s
0.916% slower
The paired mean changes were:
simg4ox: 0.887% slower, with a simple paired 95% CI of 0.396–1.379%.
GPURaytrace: 0.886% slower, with a simple paired 95% CI of 0.763–1.008%.
Head was slower in all eight GPURaytrace pairs.
Correctness
The results remained deterministic and byte-identical:
simg4ox: 1,000,000 photons, 988,940 GPU hits and 988,996 Geant4 hits in every run.
Both simg4ox hit arrays had identical SHA-256 hashes across all 16 executions.
GPURaytrace: 120,984,484 photons and 3,779,848 GPU hits in every run.
Base and head produced the same GPURaytrace hit-stream SHA-256:
GPU telemetry was collected once per second throughout the tests. The control GPU remained at 0% utilization and 8 MiB allocated, while the benchmark GPU returned to 0% and approximately 97 MiB between runs. There were no external compute processes.
Both builds reported Philox, so this measures the PR’s additional RNG-position restoration work, not XORWOW versus Philox.
Based on this uncontended rerun, the earlier apparent ~4% improvement was caused by GPU contention and should not be used as evidence for this PR. The current evidence shows preserved output with a repeatable performance regression of approximately 1%.
I reran the DUNE benchmarks after confirming that both RTX 4090s were idle. The uncontended results do not reproduce the previously reported speedup; instead, the PR is consistently about 1% slower.
Setup
* Base: `origin/main` at `681f09409`
* Head: `bc3cf5901`
* Release builds
* Fixed seed: `42`
* DUNE geometry: `dune_mock_wls_detector_box.gdml`
* Eight runs per revision:
* Five `base → head` pairs
* Three reverse-order `head → base` pairs to control for warm-up and thermal ordering
Results
Test Base median Head median Head change simg4ox, 1M photons 40.257 ms 40.719 ms 1.148% slower GPURaytrace, summed 122 GPU slices 23.8660 s 24.0880 s 0.930% slower GPURaytrace, complete simulate() 24.2023 s 24.4241 s 0.916% slower
The paired mean changes were:
* `simg4ox`: 0.887% slower, with a simple paired 95% CI of 0.396–1.379%.
* `GPURaytrace`: 0.886% slower, with a simple paired 95% CI of 0.763–1.008%.
* Head was slower in all eight `GPURaytrace` pairs.
Correctness
The results remained deterministic and byte-identical:
* `simg4ox`: 1,000,000 photons, 988,940 GPU hits and 988,996 Geant4 hits in every run.
* Both `simg4ox` hit arrays had identical SHA-256 hashes across all 16 executions.
* `GPURaytrace`: 120,984,484 photons and 3,779,848 GPU hits in every run.
* Base and head produced the same GPURaytrace hit-stream SHA-256:
`0ec0d6ec41ffbe574a1faed7cc1326bb19400321160c5ae64c99648d1af97b34`
GPU telemetry was collected once per second throughout the tests. The control GPU remained at 0% utilization and 8 MiB allocated, while the benchmark GPU returned to 0% and approximately 97 MiB between runs. There were no external compute processes.
Both builds reported Philox, so this measures the PR’s additional RNG-position restoration work, not XORWOW versus Philox.
Based on this uncontended rerun, the earlier apparent ~4% improvement was caused by GPU contention and should not be used as evidence for this PR. The current evidence shows preserved output with a repeatable performance regression of approximately 1%.
The current DUNE config that was used for this test is not a proper one. Most GPU lanes sit idle due to lifetime divergence. This PR is aimed to speed up optimized processes where performance actually matter. I will open a PR with SER that optimizes the DUNE workload. Once that is in, this PR will speed up DUNE case too then.
Yes, putting this behind a compile definition would be a good way to introduce this experimental functionality while keeping the default behavior unchanged.
Yes, putting this behind a compile definition would be a good way to introduce this experimental functionality while keeping the default behavior unchanged.
Exercise SIMPHONY_RNG_REBUILD in the GPU image matrix by compiling isolated OFF and ON PTX variants and running a deterministic five-event simg4ox workload.
Report and validate per-event photon counts, normalize multi-event hit summaries by event ID for serial and multithreaded logs, and require matching photon counts, per-event and total hit counts, and byte-identical s_hits.npy data. Also assert that the PTX artifacts differ so CI confirms both compile-time paths were built.
Add SIMPHONY_RNG_REBUILD to the CMake build-options guide with its default, PTX scope, configuration command, and fixed-seed behavior expectations.
Link the performance guide to the option and clarify that benchmarking each variant requires reconfiguring and rebuilding the CSGOptiX PTX target.
The reason will be displayed to describe this comment to others. Learn more.
The Philox overload of skipahead(n) counts scalar elements, not four-element blocks. Its implementation applies n & 3 to STATE, advances the counter by n / 4, handles the intra-block carry, and regenerates output. Therefore 4 * ctr.x + STATE is the scalar position, and reinitializing followed by skipahead(delta) restores arbitrary intra-block positions. The fixed-seed ON/OFF CTest also verifies identical per-event counts and byte-identical hit output.
Require each RNG rebuild simg4ox run to use its configured PTX artifact instead of silently falling back to the normal build output.
Replace the inverted compare_files test with an explicit comparison that passes only for two readable, distinct PTX files and preserves missing-file or comparison errors as failures.
Align PR description with opt-in RNG rebuild behavior
CSGOptiX/CMakeLists.txt:9
The PR description says RNG rebuilding is unconditional, but SIMPHONY_RNG_REBUILD defaults to OFF and guards the kernel changes. Default builds therefore retain the original RNG-state handling. Please align the description with the documented experimental, opt-in behavior and specify -DSIMPHONY_RNG_REBUILD=ON to enable it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Reduce the Philox RNG state kept live across
optixTrace()calls in thephoton-propagation loop.
Before tracing, the kernel records the current random-stream position as a
32-bit value:
After tracing, it reconstructs the 64-byte Philox state from the seed, event
index, and photon index, then uses
skipahead()to resume at the saved drawposition.
Philox is Simphony's sole device RNG, so this behavior is unconditional.
Motivation
Values that remain live across
optixTrace()contribute to the OptiXcontinuation-state footprint. Carrying the complete 64-byte Philox state can
therefore increase stack spills and L2-memory traffic.
Keeping only a 4-byte stream position across the trace trades inexpensive
Philox reconstruction for a smaller live set in the propagation kernel.
Correctness
The reconstructed state resumes at the same scalar draw position because:
ctr.xandSTATEidentify the current position within the generatedfour-value block;
optixTrace()does not consume random values; andcached normal-distribution state.
The 32-bit position delta assumes fewer than
2^32random draws per photon,which is far above the configured propagation limits.
Fixed-seed validation produced byte-identical hit output between the base and
optimized kernels.
Performance
On the reported DUNE-scale benchmark—125 million photons on an RTX 4090—the
kernel wall time improved by approximately 4%.
Additional pfRICH measurements on an RTX 4090 showed approximately 1.4–1.8%
lower propagation-kernel time, indicating that the exact improvement depends
on geometry and workload.