Repository navigation
perf(server): a returning client catches up in batches, not one event per round trip - #236
Merged
Merged
Conversation
… per round trip The thread catch-up replay taps every event to notice a delete or archive, and an effectful tap re-emits one event per stream chunk. The RPC server writes each chunk as one WebSocket frame and waits for the client's ack before the next, so a client that missed 500 events paid 500 network round trips before the thread was current. Re-chunk the replay into batches of 128. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Thread transfer impact✅ Thread transfer remains within every enforced ceiling.
Baseline: unavailable · PR result: Scenario and decoded snapshot size10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.
Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem. When a window comes back and resumes a thread, the server replays the events it missed. The replay pipeline taps each event to notice a delete or archive. An effectful
Stream.tapre-emits one event per chunk, and the RPC server writes each chunk as its own WebSocket frame, then waits for the client's ack before writing the next. Catch-up time therefore scaled as missed events × round trip.Measured on expbkt3 from the server host: resuming 500 events behind arrived as 493 frames, one per event. From Tushar's laptop in India to ap-south-1, that is 500 serial round trips before the thread is current, plus the client's per-frame handling. The server's own work is milliseconds.
Measured cost before the fix (expbkt3, a real 560 KB thread; test client holds each ack for a simulated round trip):
With 128-event batches, 500 events are 4 frames, so the same catch-up is about 4 round trips plus transfer.
Fix. Re-chunk the catch-up replay into batches of 128 (
threadReplayBatches.expbkt3.ts, one marked line inws.ts). The replay is a bounded read that already arrives page by page, so filling a batch adds no latency. The client already applies multi-item chunks as one batch (applyItemsinclient-runtime/src/state/threads.ts), so 500 missed events become 4 frames. The live stream is untouched; its coalescer already batches.Tests.
threadReplayBatches.expbkt3.test.tsreproduces the real pipeline shape (paginated read → tap → filter → map). Unbatched, it reaches the writer one event per frame; batched, it arrives in 128-event frames, in order, and a short replay is flushed at once.subscribeThreadserver tests pass (5 of 5).This is generally useful, so it's an upstream candidate.
🤖 Generated with Claude Code