Repository navigation
Fix race condition in table initialization with multiple threads - #243
Conversation
This caused wrong encoding in the OpenEXR integration, which then resulted in "Error decoding a codeblock" while decoding. Use std::call_once to avoid this. Signed-off-by: Brecht Van Lommel <brecht@blender.org>
|
This issue was found integrating the new OpenEXR HTJ2K encoding in Blender, and I confirmed that a randomly failing test now reliably passes. The bug was found and fixed by Gemini CLI, but I did review the fix carefully. A program to reproduce the issue using thread sanitizer was also generated, I've attached that. The fix resolves all thread sanitizer warnings in that program. I didn't see a way to reproduce this with existing programs included with OpenJPH. Potentially related issues: |
|
Dear Brecht, Thank you for this important discovery and for the PR. I am well with the suggested changes, but I am allergic to opaque C++ code. I am happy to do the changes if that is alright with you. Thanx again. Kind regards, |
|
Thanks for the quickly reply. Feel free to make changes. |
|
I wanted to make changes, but I am short on time. |
OpenJPH < 0.27 lazily initialized its block encoder tables with an unsynchronized static flag (fixed upstream in aous72/OpenJPH#243). When OpenEXR compresses (or decompresses) several HTJ2K chunks in parallel, the first operations in a process can race and silently write corrupt codeblocks. This made openexr-compression flaky, with intermittent failures in our CI for the "VFX2026" job (OpenEXR 3.4.15 + OpenJPH 0.24.5). Like I said, it's fixed in newer OpenJPH, but since OIIO still might encounter an older one, we need a workaround. When opening HTJ2K/LJ2K files for reading or writing, first encode one tiny single-chunk image on the calling thread (once per process) so the tables are initialized before any parallel work. OpenEXR records the OpenJPH version it was built against, so this is omitted when that is 0.27 or newer. A standalone repro with OpenEXR 3.4.15 + OpenJPH 0.24.5 wrote corrupt multi-chunk files in 54/2000 fresh processes without the priming, 0/2000 with it. Assisted-by: Claude Code / Claude Opus 5.5 Signed-off-by: Larry Gritz <lg@larrygritz.com>
…demySoftwareFoundation#5520) OpenJPH < 0.27 lazily initialized its block encoder tables with an unsynchronized static flag (fixed upstream in aous72/OpenJPH#243). When OpenEXR compresses (or decompresses) several HTJ2K chunks in parallel, the first operations in a process can race and silently write corrupt codeblocks. This made openexr-compression flaky, with intermittent failures in our CI for the "VFX2026" job (OpenEXR 3.4.15 + OpenJPH 0.24.5). Like I said, it's fixed in newer OpenJPH, but since OIIO still might encounter an older one, we need a workaround. When opening HTJ2K/LJ2K files for reading or writing, first encode one tiny single-chunk image on the calling thread (once per process) so the tables are initialized before any parallel work. OpenEXR records the OpenJPH version it was built against, so this is omitted when that is 0.27 or newer. A standalone repro with OpenEXR 3.4.15 + OpenJPH 0.24.5 wrote corrupt multi-chunk files in 54/2000 fresh processes without the priming, 0/2000 with it. Assisted-by: Claude Code / Claude Opus 5.5 Signed-off-by: Larry Gritz <lg@larrygritz.com>
This caused intermittent wrong encoding in the OpenEXR integration, which then resulted in "Error decoding a codeblock" while decoding.
Use std::call_once to resolve this.