Skip to content

perf(server): skip out-of-window usage records before aggregating - #35

Merged
andrewcai8 merged 2 commits into
mainfrom
perf/usage-scan-aggregation
Sep 22, 2026
Merged

andrewcai8 merged 2 commits into
mainfrom
perf/usage-scan-aggregation

Conversation

@andrewcai8

Copy link
Copy Markdown
Owner

The Usage page felt slow partly because every getUsageSummary walked every cached transcript record, about 275k on a real machine, even for the past-24h view, which uses about 49k. It also rebuilt the Codex dedupe key (Schema.encodeSync) for each record on every request, and stat'ed about 3,800 transcript files one at a time.

How it's fixed:

  • The aggregator exposes a cheap admits(timestamp) check. Records outside the window are skipped before any key is built: exactly for hourly windows, and with one day of slack for day windows so any time-zone offset is safe. The exact day check still decides. Keyed Claude/Grok records still go through add, so a copy seen outside the window still suppresses a later copy inside it, as before.
  • The per-file loop moved into addTranscript, which builds each Codex key once per cached record and reuses it.
  • listTranscriptFiles stats files 32 at a time and keeps walk order, which decides which copy of a duplicate wins.

Results are unchanged. An equality test compares addTranscript with the old loop across five windows (hourly LA, day LA, +14h, −11h, a UTC month), including moved Codex rollouts and a Claude key first seen out of window. A differential run on the real 275k-record cache matched in 25 of 25 zone/window combinations.

Measured on a warm cache (e2e UsageService.readSummary without Cursor, two runs each):

Window Before After
24h 430 / 498 ms 154 / 212 ms
7d 566 / 705 ms 272 / 284 ms
30d 642 / 668 ms 392 / 475 ms
90d 1071 / 619 ms 453 / 484 ms

The machine was under heavy load during these runs, so the numbers are noisy.

🤖 Generated with Claude Code

andrewcai8 and others added 2 commits September 22, 2026 19:19
Every usage summary, even the 24h window, walked all ~275k cached records
and rebuilt each Codex dedupe key with a Schema JSON encode per request.

The aggregator now exposes a cheap instant bound (exact for hourly
windows, a one-day superset for zoned day windows), and unkeyed records
outside it are skipped before keying. Equal Codex identities share a
timestamp, so skipping never shifts an in-window occurrence count, and
keyed records still register their key out of window. Codex identities
are memoised per record.

Summaries are identical to the old loop on the real 275k-record cache
across five zones and five windows. Warm 24h summary: 430-498ms to
154-212ms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Listing ~3,900 transcripts awaited one stat at a time. Stats now run 32
at a time while results keep walk order, which decides which copy of a
duplicated record wins. Warm listing drops from ~60-110ms to ~35ms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Sep 22, 2026
@andrewcai8

Copy link
Copy Markdown
Owner Author

Verdict: PASS+NOTES

Head: 6667d6f (independent verifier, fresh worktree)

What I ran

  • vp test run apps/server/src/usage: 8 files, 92 tests pass.
  • apps/server typecheck (tsc --noEmit): no errors.
  • /tmp/usage-perf-agg/differential.ts re-pointed at this worktree. Its ref/usageAggregation.ts matches main apart from import paths. All 25 zone/window combinations are identical against the real 275k-record cache.
  • A throwaway fuzz differential, deleted afterward: 3000 random windows, of which 2732 had non-empty output, comparing addTranscript with main's loop on buckets and session ids. Zones: NY/London/Lord_Howe DST transition days, Kiritimati, Etc/GMT±12/14, Chatham, Kathmandu, St_Johns, Apia, plus a year boundary. Timestamps sit at ±1 ms around hourly bounds and local/UTC midnights. The mix includes empty sessionIds, codex records that carry a parser dedupeKey, Claude/Grok keys reused across files (including copies seen out of window first), records shared by reference between files (the same situation as a record in both the live listing and the retained fileCache), and a WeakMap warm-up pass on a different window before each comparison. No mismatches.

Findings

  • The admits bounds hold. The day window covers [sinceDay 00:00Z - 1d, untilDay 00:00Z + 2d), which contains every local day for offsets from -12h to +14h. The hourly check is exact, as before.
  • Codex key:occurrence is unchanged. The identity includes timestampMs, so the prefilter drops all copies of an identity or none of them. The WeakMap value does not depend on the window, and records are spread rather than mutated.
  • The concurrent stat keeps walk order because each result goes into its own index slot.
  • Note: duplicatesDropped and outOfWindow from finish() now undercount, because prefiltered records never reach add. Nothing reads them (UsageService uses only buckets), so this is harmless. Either delete them or update the outOfWindow doc.

@andrewcai8
andrewcai8 enabled auto-merge (squash) September 22, 2026 23:30
@andrewcai8
andrewcai8 merged commit 3c75242 into main Sep 22, 2026
17 checks passed
@github-actions

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire 13.5 KiB 13.5 KiB −48 B (−0.3%) 15.1 KiB ✅
Codex Thread snapshot wire 7.1 KiB 7.1 KiB +1 B (+0.0%) 7.3 KiB ✅
Codex Live turn WebSocket wire 6.5 KiB 6.4 KiB −49 B (−0.7%) 7.8 KiB ✅
Codex Live turn WebSocket decoded 56.3 KiB 56.2 KiB −88 B (−0.2%) 66.4 KiB ✅
Codex Live turn messages 10 8 −2 (−20.0%) 21 ✅
Claude Total thread wire 13.5 KiB 13.5 KiB +7 B (+0.1%) 15.1 KiB ✅
Claude Thread snapshot wire 7.1 KiB 7.1 KiB +1 B (+0.0%) 7.3 KiB ✅
Claude Live turn WebSocket wire 6.4 KiB 6.4 KiB +6 B (+0.1%) 7.8 KiB ✅
Claude Live turn WebSocket decoded 57.0 KiB 57.0 KiB 0 B (0.0%) 66.4 KiB ✅
Claude Live turn messages 9 9 0 (0.0%) 21 ✅

Baseline: bf8fd20 · PR result: 6667d6f · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 113.9 KiB
  • Claude decoded thread snapshot: 114.6 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@andrewcai8
andrewcai8 deleted the perf/usage-scan-aggregation branch October 4, 2026 04:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant