Skip to content

Promote expbkmain: bkt3 performance work (PRs #214–#239) - #224

Merged
bk-agent-01 merged 70 commits into
bkmainfrom
expbkmain
Sep 27, 2026
Merged

bk-agent-01 merged 70 commits into
bkmainfrom
expbkmain

Conversation

@tusharbhardwaj-bk

@tusharbhardwaj-bk tusharbhardwaj-bk commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Do not merge without warning the team. A bkmain deploy restarts t3-bkmain.service and ends every live session on bkt3. The human merges this.

What this promotes

36 commits from PRs #214–#223 and #225–#234, all merged to expbkmain and verified on expbkt3. Nothing else: git log --no-merges origin/bkmain..origin/expbkmain shows only these, and bkmain has no commits that expbkmain lacks. The 9 commits by other authors are the upstream cherry-picks in #222.

Server: fewer statements, less event-loop blocking

PR Change Measured
#214 Drop duplicate provider status events 10 identical pings → 1 event; real turn on expbkt3: 0 duplicates, against 58% before (77% in the prod audit)
#215 Startup reconcile reads only waiting threads; per-phase startup timing logs 28.7 s cold / 0.9 s warm → 0.4 ms on a prod copy; expbkt3 boots to Listening in 3–4 s
#227 Execution overlay batched 588 → 3 statements per shell read; 61 → 22 ms on a prod copy
#228 Shell resume filters visibility once ~1,576 → ~11 statements per team-mode reconnect (with #227)
#229 Reaper decodes only live bindings 1,355 → 3 rows; 37–67 ms → 1–4 ms (prod p90 was 1.86 s); expbkt3 bindingsMs: 1
#219 Thread cost reads only its own transcripts; 23 MB cache write debounced per-thread file listing; cache rewrite at most every 5 min
#218 PRAGMA optimize every 6 h (never at startup: the first run is 2.8 s) plans unchanged for the 7 hot queries
#234 Summary reactors skip activity payloads 58–206 → 33–40 ms per load (feature is off in prod)
#222 9 upstream cherry-picks no provider refresh per connect, git process cap, bounded provider logs, cheaper checkpoints, etc.

New fork APIs for the Linear bridge (switched off until the gated steps below)

Client (reaches desktop users only with the next desktop build)

CI and deploy after the repository move to iamtushar324/bkt3code

Verified on expbkt3 (ed3c051fd, deployed 2026-09-27 07:01 UTC)

  • Deploy timer: installs new builds automatically again (it deferred correctly while validation ran).
  • Startup: 3 s to Listening, with every phase timed in the journal.
  • Event feed: paging, limit, wait, 410, 400 and 401 all correct.
  • PR-state: no match, bad body, and a replay ignored as unchanged all correct.
  • 3-minute stress run:
    • load: 40 long-polling followers, 15 commands/s, and 10 concurrent PR-state writers;
    • results: 99,031 feed pages with 0 gaps, 0 duplicates, 0 ordering errors and 0 failed requests; 1,912 of 1,912 PR-state writes returned 200; 511 of 511 commands succeeded;
    • memory returns to baseline after the load (412 MB after ~2 min idle; not a leak).
  • Real Claude turn: completed in 15 s, with 0 duplicate session events.

After merging (human steps)

  1. bkt3 won't auto-deploy this merge. bkmain's current auto-deploy.sh still has the redirect bug that fix(deploy): auto-deploy follows repository redirects #223 fixes. Once the deploy-bkt3.yml push run for the merge commit is green, run one manual deploy from a human shell or the expbkt3 instance, never from a bkt3 session:

    T3_DEPLOY_GH_TOKEN_FILE=/home/ubuntu/.config/t3-deploy/gh-token \
      /home/ubuntu/.local/bin/t3-deploy-with-github-token \
      /home/ubuntu/repos/t3code-bkmain/deploy/bkt3/deploy.sh <merge-sha> <run-id>

    The same deploy fixes the t3.dev timer, whose script runs from this checkout.

  2. Biggest remaining prod win, gated:

    • give the bridge a token with external-sync:write (docs/operations/external-pr-sync.md);
    • deploy the bridge's event-feed follower and PR-state sender (session B);
    • then set T3_EXTERNAL_PR_SYNC=1 in deploy/bkt3/start.sh.

    That removes the per-minute PR sweep (92% of git spawns on bkt3) and the bridge's shell polling.

  3. Cut a desktop build, so desktop users get perf(web): window focus probes instead of resyncing; desktop wake reconnects at once #220, perf: cherry-pick upstream connection, git and checkpoint fixes #222 (keep-alive), fix(client-runtime): a busy thread stream no longer runs out of retries #230, perf(client-runtime): context-window updates stop re-sorting a thread's history #231 and perf(web): history deepening waits for a quiet, visible thread #232.

Model/harness: Claude Opus 5.5 via Claude Code in T3 Code.

Added 2026-09-27: reconnect and catch-up lag (#236–#239)

Tushar reported that returning to the desktop app and syncing new messages still wasn't instant. Measured on expbkt3 and fixed:

PR What was slow Before After
#236 Catch-up replay went one event per WebSocket frame, each waiting for the client's ack 500 missed events: 17.9 s at a 35 ms round trip (493 frames) 0.3 s (6 frames)
#237 Desktop window open waited on the bundled local backend before connecting to the primary ~10.5 s gap on 2 of 3 opens in the traces at most 1.5 s
#238 One return to the window resynced twice (focus + visibility), and a return after 60 s away only reached the first subscriber shell and 4 threads resubscribed twice, 2 s apart once, for every subscriber
#239 An execution frame batched with events was dropped, which could leave a chat stuck on "Running" dropped applied

This promotion also carries #235 (worktree setup progress row). The desktop fixes (#237, #238, #239) reach users through a desktop build; staging build v0.0.41-staging-nightly.20260927.8 has them.

🤖 Generated with Claude Code

bk-agent-01 and others added 30 commits September 26, 2026 08:22
Claude reports a status ping on every model request inside a turn, and
ingestion turned each one into a thread.session.set whose only change was
updatedAt (2,040 of 2,666 session-set events in a 3.6 h prod sample).

A fork helper now skips the dispatch for session.state.changed when status,
activeTurnId, lastError, providerThreadId, runtimeMode, providerName and
providerInstanceId all equal the projected session. Lifecycle events always
dispatch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The waiting-intent repair in reconcileStartup read every approval and
user-input activity in projection_thread_activities (954k rows in prod) on
every boot. It now reads only the threads that have a waiting intent, which
keeps it on the thread_id index: 28.7 s cold / 0.9 s warm -> 0.4 ms on a
copy of the prod database.

Each startup phase and the durable reconcile now log their duration, so the
next slow boot names its phase in the journal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
GET /api/orchestration/events?after=<seq>&limit=<1..1000>&wait=<0..25>
returns events after a cursor in sequence order, shaped like the WebSocket
replay. An empty read is held until the engine publishes a newer event or
the wait elapses (woken by the live event stream, no polling). A cursor
ahead of the log gets 410 cursor-invalid.

This lets the Linear bridge stop polling the full shell every 2 s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
POST /api/orchestration/pull-request-state lets the Linear bridge push PR
state built from GitHub webhooks. bkt3 makes no outside calls: it refreshes
the snapshot of threads linked to the PR and sets the branch PR of
unarchived threads on its head branch in the same repository, with the same
commands the upstream PR reactors dispatch. Older or identical writes are
ignored; command ids derive from the delivery id.

The endpoint needs the new external-sync:write scope, granted with
`t3 auth session issue --with-scope external-sync:write`.

T3_EXTERNAL_PR_SYNC=1 stops ThreadPullRequestReactor and
PullRequestSyncReactor from starting. It is not set anywhere yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Prod has never been analysed (no sqlite_stat1), so the planner chooses among
overlapping indexes without statistics. Every 6 h the server now runs
PRAGMA analysis_limit=400; PRAGMA optimize=0x10002.

It never runs at startup: the first-ever run took 2.8 s on a copy of the
prod database, and node:sqlite blocks the event loop while it runs. Later
runs take ~0 ms. T3_SQLITE_OPTIMIZE=0 turns it off.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…edupe

perf(server): drop duplicate provider session status events
threadUsage.get (the chat header's cost chip) walked and parsed every
transcript modified since the thread started, then rewrote the whole 23 MB
usage-scan-cache.json on every call (7.6 s average, 21 s max in prod).

It now lists only the thread's own files (Claude <sessionId>.jsonl plus its
subagents folder, Codex rollout-*-<sessionId>.jsonl), and a thread read
rewrites the scan cache at most every 5 minutes; the service flushes the
cache on shutdown.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…onnects at once

Every window focus fired "application-active", so each alt-tab rebuilt the
shell subscription and replayed every open thread. Focus now emits a new
probe-only "application-focus" wakeup unless the window was blurred or
hidden for at least 60 s.

After sleep the desktop window stays visible, so the focus probe waited out
its 15 s timeout on a dead socket. The desktop shell now forwards the OS
resume and unlock-screen events to the renderer, which reconnects at once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The check diffs HEAD against its merge-base with upstream, so a
`git cherry-pick -x` of an upstream fix the fork has not merged yet showed up
as unmarked fork edits and failed the required check. Wrapping cherry-picked
code in markers would create conflicts in the very merge that brings the same
commit in.

A line now only counts as a fork edit when it also differs from the upstream
tip (FORK_UPSTREAM_REF, which CI already sets to upstream/main).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ingdotgg#11811)

Co-authored-by: Bil0000 <bilal.bakr.elsherif@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
(cherry picked from commit 7931227)
…ngdotgg#8309)

Co-authored-by: Julius Marminge <julius0216@outlook.com>
(cherry picked from commit a5da327)
…gg#13554)

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
(cherry picked from commit 99641fd)
The cherry-picked test from 99641fd builds an OrchestrationThread without
the fields the fork adds (owner, members, summaries, source control profile).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…cile

perf(server): startup reconcile no longer scans every activity
…state

feat(server): accept pull-request state from an external syncer
perf(server): keep sqlite query statistics fresh
perf(server): thread cost reads only that thread's transcripts
perf(web): window focus probes instead of resyncing; desktop wake reconnects at once
ci: fork-marker check ignores lines identical to upstream
perf: cherry-pick upstream connection, git and checkpoint fixes
The bkt3code repository moved owner, so api.github.com answers the old
beknown-work/bkt3code path with a 301. curl -f does not treat a redirect as
an error, jq then failed on the redirect body ("Cannot iterate over null"),
and t3-expbkt3-deploy.service exited 5 every minute, so the timer stopped
installing new builds.

auto-deploy.sh (expbkt3 and bkt3) now follows redirects, treats any response
without a workflow_runs array as "deferred" instead of failing the unit, and
authenticates with the GH_TOKEN the systemd wrapper already exports.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Same 301 failure as expbkt3 and bkt3: deploy/t3/auto-deploy.sh (run by
t3-beknown-deploy.timer from the bkmain checkout) would stop at the next
t3main push.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ility

perf(server): shell resume filters visibility once, not per thread
…bindings

perf(server): reaper sweep decodes only live provider bindings
…-budget

fix(client-runtime): a busy thread stream no longer runs out of retries
…vity-append

perf(client-runtime): context-window updates stop re-sorting a thread's history
…ening

perf(web): history deepening waits for a quiet, visible thread
…ctivities

perf(server): summary reactors stop loading activities they never read
@github-actions

github-actions Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

ℹ️ No successful main baseline artifact is available yet. This run establishes the initial measurement.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire — 13.8 KiB — 15.1 KiB ✅
Codex Thread snapshot wire — 7.3 KiB — 7.8 KiB ✅
Codex Live turn WebSocket wire — 6.5 KiB — 7.8 KiB ✅
Codex Live turn WebSocket decoded — 57.1 KiB — 66.4 KiB ✅
Codex Live turn messages — 9 — 21 ✅
Claude Total thread wire — 13.8 KiB — 15.1 KiB ✅
Claude Thread snapshot wire — 7.3 KiB — 7.8 KiB ✅
Claude Live turn WebSocket wire — 6.5 KiB — 7.8 KiB ✅
Claude Live turn WebSocket decoded — 57.9 KiB — 66.4 KiB ✅
Claude Live turn messages — 9 — 21 ✅

Baseline: unavailable · PR result: e1d692c · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 114.9 KiB
  • Claude decoded thread snapshot: 115.6 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@tusharbhardwaj-bk
tusharbhardwaj-bk deployed to bk-desktop-staging September 27, 2026 06:59 — with GitHub Actions Active
@tusharbhardwaj-bk tusharbhardwaj-bk changed the title Promote expbkmain: bkt3 performance fixes (steps 1–3), deploy redirect fix Promote expbkmain: bkt3 performance work (PRs #214–#234) Sep 27, 2026
bk-agent-01 and others added 2 commits September 27, 2026 09:04
While a new worktree thread is created and its setup script runs, the chat
showed "Working… / Thinking" as if the agent already had the prompt. The
durable bootstrap already tracks worktree, setup and agent phases on the
execution intent; the UI now uses them.

- Web: the working row becomes an expandable "Setting up workspace" checklist
  (worktree, setup script with Show output, agent start). It collapses to a
  "Workspace ready" checkmark once the agent takes over, and "Thinking" only
  appears once the agent turn is running.
- Mobile: the floating status pill names the current step and the feed
  holds back the Thinking row until the agent starts.
- Server: publish the execution snapshot when the setup step starts, so
  clients see "worktree created, setup running" live instead of only at the
  end of preparation.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
feat(web,mobile): show worktree setup progress instead of "Thinking"
@bk-agent-01
bk-agent-01 deployed to bk-desktop-staging September 27, 2026 09:36 — with GitHub Actions Active
bk-agent-01 and others added 5 commits September 27, 2026 14:41
… per round trip

The thread catch-up replay taps every event to notice a delete or archive,
and an effectful tap re-emits one event per stream chunk. The RPC server
writes each chunk as one WebSocket frame and waits for the client's ack
before the next, so a client that missed 500 events paid 500 network round
trips before the thread was current. Re-chunk the replay into batches of 128.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…o the primary by 10 s

The platform source builds every registration before handing any to the
registry, so the remote primary only connects once each desktop-local backend
has described itself. A managed build starts its bundled backend after the
window opens, and a backend still starting holds requests until it is ready,
so the descriptor fetch waited out its 10 s timeout. Cap it at 1.5 s; a ready
loopback backend answers in milliseconds and a slow one is retried on the
next 3 s poll.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t fires

One return to the window fires focus, visibilitychange and sometimes online,
and each "application-active" restarts every open subscription and throws
away the catch-up the previous one had in flight. expbkt3 traces showed the
shell and four threads resubscribed at 14:13:24 and again 2 s later. After
a resync or reconnect, any further "application-active" inside 5 s only
probes the connection.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ing again

Waking a laptop fires the OS resume and, once the password is typed,
unlock-screen, often 10-20 s apart. Each was forwarded as a reconnect, so
the unlock tore down the healthy socket the resume had just built and
replayed every thread again. A reconnect within 60 s of the previous one
now only probes; a dead socket fails the probe and reconnects at once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
perf(server): a returning client catches up in batches, not one event per round trip
@tusharbhardwaj-bk
tusharbhardwaj-bk deployed to bk-desktop-staging September 27, 2026 15:09 — with GitHub Actions Active
tusharbhardwaj-bk and others added 2 commits September 27, 2026 20:40
…ery-timeout

perf(desktop): a starting local backend no longer delays connecting to the primary by 10 s
…return-wakeups

# Conflicts:
#	apps/web/src/connection/platform.ts
@tusharbhardwaj-bk
tusharbhardwaj-bk deployed to bk-desktop-staging September 27, 2026 15:17 — with GitHub Actions Active
bk-agent-01 and others added 5 commits September 27, 2026 15:23
…d, not dropped

The batch fast path in applyItems predates the fork's execution frames and
only handled events and completion markers, so an execution frame that
arrived in the same transport batch as events was silently dropped. A chat
resumed after its turn ended could keep showing "Running". Batching the
catch-up replay makes that batch shape common.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…just the first

The focus tracker is module-wide but every wakeup subscriber registers its
own focus listener, so one focus event reached onFocus once per subscriber
and only the first saw the return; the shell and thread streams only got a
probe. Classify each focus event once by its timeStamp.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
fix(client-runtime): an execution frame batched with events is applied, not dropped
perf(web): one return to the window resyncs once, and reaches every subscriber
@tusharbhardwaj-bk
tusharbhardwaj-bk deployed to bk-desktop-staging September 27, 2026 16:04 — with GitHub Actions Active
@tusharbhardwaj-bk tusharbhardwaj-bk changed the title Promote expbkmain: bkt3 performance work (PRs #214–#234) Promote expbkmain: bkt3 performance work (PRs #214–#239) Sep 27, 2026
@bk-agent-01
bk-agent-01 merged commit aa846e7 into bkmain Sep 27, 2026
30 checks passed

This branch was successfully deployed

1 active deployment
bk-desktop-staging — e1d692cb Deployed Sep 27, 2026 by tusharbhardwaj-bk via Publish #151
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants