Skip to content

fix(openexr): Work around OpenJPH < 0.27 HTJ2K encoder init race - #5520

Merged
lgritz merged 1 commit into
AcademySoftwareFoundation:mainfrom
lgritz:lg-htj2k-race
Sep 30, 2026
Merged

lgritz merged 1 commit into
AcademySoftwareFoundation:mainfrom
lgritz:lg-htj2k-race

Conversation

@lgritz

@lgritz lgritz commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

OpenJPH < 0.27 lazily initialized its block encoder tables with an unsynchronized static flag (fixed upstream in aous72/OpenJPH#243). When OpenEXR compresses (or decompresses) several HTJ2K chunks in parallel, the first operations in a process can race and silently write corrupt codeblocks. This made openexr-compression flaky, with intermittent failures in our CI for the "VFX2026" job (OpenEXR 3.4.15 + OpenJPH 0.24.5).

Like I said, it's fixed in newer OpenJPH, but since OIIO still might encounter an older one, we need a workaround.

When opening HTJ2K/LJ2K files for reading or writing, first encode one tiny single-chunk image on the calling thread (once per process) so the tables are initialized before any parallel work. OpenEXR records the OpenJPH version it was built against, so this is omitted when that is 0.27 or newer. A standalone repro with OpenEXR 3.4.15 + OpenJPH 0.24.5 wrote corrupt multi-chunk files in 54/2000 fresh processes without the priming, 0/2000 with it.

Assisted-by: Claude Code / Claude Opus 5.5

OpenJPH < 0.27 lazily initialized its block encoder tables with an
unsynchronized static flag (fixed upstream in aous72/OpenJPH#243).
When OpenEXR compresses (or decompresses) several HTJ2K chunks in
parallel, the first operations in a process can race and silently
write corrupt codeblocks. This made openexr-compression flaky, with
intermittent failures in our CI for the "VFX2026" job (OpenEXR 3.4.15
+ OpenJPH 0.24.5).

Like I said, it's fixed in newer OpenJPH, but since OIIO still might
encounter an older one, we need a workaround.

When opening HTJ2K/LJ2K files for reading or writing, first encode one
tiny single-chunk image on the calling thread (once per process) so
the tables are initialized before any parallel work. OpenEXR records
the OpenJPH version it was built against, so this is omitted when that
is 0.27 or newer. A standalone repro with OpenEXR 3.4.15 + OpenJPH
0.24.5 wrote corrupt multi-chunk files in 54/2000 fresh processes
without the priming, 0/2000 with it.

Assisted-by: Claude Code / Claude Opus 5.5

Signed-off-by: Larry Gritz <lg@larrygritz.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

A failed warm-up is silently marked complete, leaving later parallel HTJ2K operations exposed to the race.

Review effort: Balanced
Findings: 1 High severity · 1 Low severity

Open (2)
What changed in this PR

This PR works around a cold-start race in older OpenJPH versions by priming HTJ2K encoding before OpenEXR processes chunks in parallel.

Changes:

  • Add a once-per-process warm-up for affected OpenJPH versions.
  • Invoke it when opening HTJ2K files for writing or reading.
File Description
src/​openexr.imageio/​exroutput.cpp Adds the warm-up and calls it from output opens.
src/​openexr.imageio/​exrinput.cpp Calls the warm-up from the OpenEXR reader.
src/​openexr.imageio/​exrinput_c.cpp Calls the warm-up from the core reader.
src/​openexr.imageio/​exr_pvt.h Declares the shared warm-up function.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +148 to +149
} catch (...) {
}

@lgritz lgritz Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My robot and I think that Copilot is wrong about this. Ignoring.

Comment on lines +145 to +147
Imf::OutputFile out(stream, header, 0 /* no threads */);
out.setFrameBuffer(fb);
out.writePixels(8);
@lgritz

lgritz commented Sep 30, 2026

Copy link
Copy Markdown
Collaborator Author

Need this for broken CI. Merging.

@lgritz
lgritz merged commit f4c0806 into AcademySoftwareFoundation:main Sep 30, 2026
32 checks passed
lgritz added a commit to lgritz/OpenImageIO that referenced this pull request Sep 30, 2026
…demySoftwareFoundation#5520)

OpenJPH < 0.27 lazily initialized its block encoder tables with an
unsynchronized static flag (fixed upstream in aous72/OpenJPH#243). When
OpenEXR compresses (or decompresses) several HTJ2K chunks in parallel,
the first operations in a process can race and silently write corrupt
codeblocks. This made openexr-compression flaky, with intermittent
failures in our CI for the "VFX2026" job (OpenEXR 3.4.15 + OpenJPH
0.24.5).

Like I said, it's fixed in newer OpenJPH, but since OIIO still might
encounter an older one, we need a workaround.

When opening HTJ2K/LJ2K files for reading or writing, first encode one
tiny single-chunk image on the calling thread (once per process) so the
tables are initialized before any parallel work. OpenEXR records the
OpenJPH version it was built against, so this is omitted when that is
0.27 or newer. A standalone repro with OpenEXR 3.4.15 + OpenJPH 0.24.5
wrote corrupt multi-chunk files in 54/2000 fresh processes without the
priming, 0/2000 with it.

Assisted-by: Claude Code / Claude Opus 5.5

Signed-off-by: Larry Gritz <lg@larrygritz.com>
@lgritz
lgritz deleted the lg-htj2k-race branch October 1, 2026 01:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants