perf: compress undo snapshots with LZ4 level 1 instead of 5 - #40
Merged
Merged
Conversation
The interactions tensor is overwhelmingly zeros, so LZ4 level 1 produces the same compressed size as level 5 (0.4 MB from a 2.9 GB 8x694x512x512 fp16 tensor, measured with the exact cparams of _blosc2_cparams) while cutting the compress from 271 ms to 80 ms on 8 threads. The post-predict snapshot is submitted asynchronously and deliberately overlaps user think-time, but a fast follow-up interaction lands inside the compression burst and contends with request handling; in our deployment (nnInteractive behind a FastAPI server) that showed up as an intermittent ~330 ms added latency on rapid consecutive prompts, measured from the client as time-to-first-byte minus server compute. A ~3.4x shorter burst shrinks that window proportionally. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Member
|
Let's give that a try, thanks Joey! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
One-line change to
_blosc2_cparams:clevel5 → 1 (plus a comment recording why).Why
The undo snapshot compresses the full interactions tensor after every prediction. That tensor is overwhelmingly zeros, so LZ4's search depth buys nothing:
(Benchmarked with the exact cparams this method returns — LZ4, NOFILTER — on a tensor with a realistic nonzero region; sizes were byte-identical.)
The snapshot is submitted asynchronously and deliberately overlaps user think-time, which works well — until a fast follow-up interaction lands inside the compression burst and contends with it. In our deployment (nnInteractive sessions behind a FastAPI server, multiple users) this showed up as an intermittent ~330 ms of added latency on rapid consecutive prompts, measured from the client as time-to-first-byte minus server-side compute. A ~3.4× shorter burst shrinks that collision window proportionally, at zero cost in snapshot size or undo behavior.
Related finding (not in this diff)
If you ever consider deprioritizing the snapshot thread instead: nice-then-restore does not survive containers — Docker's default seccomp/cap profile drops
CAP_SYS_NICE, so restoring priority fails withEPERMeven as in-container root. A permanently self-niced dedicated thread works (that's what we run alongside this change), but it's deployment-sensitive, so we're only upstreaming the universally-safe clevel change.🤖 Generated with Claude Code