Skip to content

perf: compress undo snapshots with LZ4 level 1 instead of 5 - #40

Merged
MIC-DKFZ-Bot merged 1 commit into
MIC-DKFZ:masterfrom
Joeycho:perf/snapshot-lz4-level-1
Sep 10, 2026
Merged

MIC-DKFZ-Bot merged 1 commit into
MIC-DKFZ:masterfrom
Joeycho:perf/snapshot-lz4-level-1

Conversation

@Joeycho

@Joeycho Joeycho commented Aug 20, 2026

Copy link
Copy Markdown
Contributor

What

One-line change to _blosc2_cparams: clevel 5 → 1 (plus a comment recording why).

Why

The undo snapshot compresses the full interactions tensor after every prediction. That tensor is overwhelmingly zeros, so LZ4's search depth buys nothing:

clevel compress time (2.9 GB, 8×694×512×512 fp16, 8 threads) compressed size
5 (current) 271 ms 0.4 MB
1 (this PR) 80 ms 0.4 MB

(Benchmarked with the exact cparams this method returns — LZ4, NOFILTER — on a tensor with a realistic nonzero region; sizes were byte-identical.)

The snapshot is submitted asynchronously and deliberately overlaps user think-time, which works well — until a fast follow-up interaction lands inside the compression burst and contends with it. In our deployment (nnInteractive sessions behind a FastAPI server, multiple users) this showed up as an intermittent ~330 ms of added latency on rapid consecutive prompts, measured from the client as time-to-first-byte minus server-side compute. A ~3.4× shorter burst shrinks that collision window proportionally, at zero cost in snapshot size or undo behavior.

Related finding (not in this diff)

If you ever consider deprioritizing the snapshot thread instead: nice-then-restore does not survive containers — Docker's default seccomp/cap profile drops CAP_SYS_NICE, so restoring priority fails with EPERM even as in-container root. A permanently self-niced dedicated thread works (that's what we run alongside this change), but it's deployment-sensitive, so we're only upstreaming the universally-safe clevel change.

🤖 Generated with Claude Code

The interactions tensor is overwhelmingly zeros, so LZ4 level 1 produces
the same compressed size as level 5 (0.4 MB from a 2.9 GB 8x694x512x512
fp16 tensor, measured with the exact cparams of _blosc2_cparams) while
cutting the compress from 271 ms to 80 ms on 8 threads.

The post-predict snapshot is submitted asynchronously and deliberately
overlaps user think-time, but a fast follow-up interaction lands inside
the compression burst and contends with request handling; in our
deployment (nnInteractive behind a FastAPI server) that showed up as an
intermittent ~330 ms added latency on rapid consecutive prompts,
measured from the client as time-to-first-byte minus server compute.
A ~3.4x shorter burst shrinks that window proportionally.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@FabianIsensee

Copy link
Copy Markdown
Member

Let's give that a try, thanks Joey!

@MIC-DKFZ-Bot
MIC-DKFZ-Bot merged commit 04a01f8 into MIC-DKFZ:master Sep 10, 2026
@Joeycho
Joeycho deleted the perf/snapshot-lz4-level-1 branch September 10, 2026 12:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants