Skip to content

[Bug]: Assistant streaming causes full-thread projection scans for every text delta #5719

Description

@cheruvian

Steps to reproduce

  1. Run T3 Code 0.0.32 with a mature thread containing substantial message and activity history.
  2. Enable Settings → Assistant output → Stream assistant messages.
  3. Start a Codex turn that streams a moderately long response.
  4. Observe CPU and RSS for the T3 server process (not only the provider subprocess).
  5. Compare with assistant streaming disabled.

The degradation grows with both the number of streamed response chunks and the amount of existing thread history.

Expected behavior

Streaming an assistant response should have approximately constant per-delta projection cost. Assistant text deltas should not repeatedly load and decode complete message, proposed-plan, activity, and approval collections.

Actual behavior

Each provider assistant-text delta is dispatched as a durable thread.message.assistant.delta command and becomes a thread.message-sent event.

Every event is then applied serially to all nine projectors. The threads projector handles every thread.message-sent event by invoking refreshThreadShellSummary, which reloads all messages, proposed plans, activities, and pending approvals for that thread.

Sanitized measurements from a real 0.0.32 database:

  • 3,610 logical messages produced 128,972 thread.message-sent events
  • Average: 35.7 persisted chunk events per logical message
  • Maximum observed: 2,436 events for one 11,646-character message
  • Busiest thread: 25,878 orchestration events
  • With approximately one new orchestration event per second, the T3 server sustained 110–165% CPU
  • Server RSS oscillated between approximately 1.4 GB and 2.6 GB

The observed cost therefore scales with streamed chunks × accumulated thread history. Repeatedly loading and decoding the full activity collection also creates significant allocation and garbage-collection pressure.

Relevant source paths

  • apps/server/src/orchestration/Layers/ProviderRuntimeIngestion.ts: when enableAssistantStreaming is true, each assistant delta is individually dispatched (currently around lines 1638–1683).
  • apps/server/src/orchestration/Layers/ProjectionPipeline.ts: refreshThreadShellSummary loads all four thread collections (currently around lines 547–589).
  • The same projector calls that refresh for every thread.message-sent event (currently around lines 845–860).
  • Every event runs nine projectors serially, with a transactional projector cursor update for each one (currently around lines 1601–1692).

This appears to undermine the intent of #1647, which removed an earlier thread-wide scan from message projection.

Issue #2761 is adjacent but appears distinct: that report concerns large WebSocket snapshots and reconnect behavior, while this report concerns synchronous persistence/projection work during active streaming.

Suggested direction

The narrowest fix may be to avoid refreshThreadShellSummary for assistant thread.message-sent events. Assistant text cannot change latestUserMessageAt, pending approvals, pending user input, or actionable proposed-plan state.

More generally, shell-summary fields could be maintained incrementally or through targeted aggregate queries for the event types that can actually affect each field. A useful regression contract would be: projecting an assistant delta performs a constant number of queries independent of thread history and does not query activities, proposed plans, or approvals.

Impact

Major degradation or frequent failure.

Version or commit

T3 Code Alpha 0.0.32.

Environment

macOS on Apple Silicon, desktop app, Codex app-server provider, assistant streaming enabled.

Logs or stack traces

No raw logs or database artifacts are attached because they can contain private agent content. The measurements above contain aggregate counts only.

Workaround

Disable Stream assistant messages. The server checks this setting while processing deltas, so subsequent text is handled by the existing buffered path without terminating the provider session. Starting a fresh thread reduces the history-dependent portion of the cost but does not address the per-delta transaction amplification.

Activity

  1. cheruvian commented on Aug 8, 2026

    @cheruvian
    Author

    One scope clarification from follow-up profiling: disabling assistant streaming removes the largest source of thread.message-sent deltas, but it does not eliminate the same projection amplification for activity events.

    thread.activity-appended follows the same shell-summary refresh path, which reloads the thread-wide message/plan/activity/approval inputs. In a sanitized 10-second sample after streaming was disabled, 7 new activity events corresponded with 55 projector runs, 7 shell-summary refreshes, and 62 SQL transaction traces while server CPU remained elevated.

    So the fix should cover both high-frequency assistant deltas and activity-appended events, ideally by applying incremental projection updates or coalescing shell refreshes. Closing the renderer eliminated concurrent VCS work but did not eliminate this projection load, helping separate the two issues.

  2. dain commented on Aug 12, 2026

    @dain

    I observed a cross-thread consequence of this same projection amplification: a high-activity thread can prevent unrelated newly submitted turns from showing lifecycle or assistant-output events for several minutes.

    Sanitized evidence from a current main checkout:

    • A new turn was requested at 20:02:29.
    • Its provider lifecycle event was timestamped 20:02:38, indicating that the provider started promptly.
    • That event was not appended until after a client event from 20:10:18.
    • There were 647 orchestration events ahead of it, including 562 thread.activity-appended events.
    • One unrelated thread contributed 526 of the queued events.
    • The incoming burst totaled only approximately 3.74 MB.

    The expensive thread had accumulated:

    • 27,303 activity rows / 131.2 MB of activity payloads
    • 23,778 message rows / 9 MB of message text

    Its 526 queued events consisted of 465 routine tool/context activity events and 61 assistant-message events. None needed a full shell-summary rebuild, but each triggered refreshThreadShellSummary.

    A live trace from the same running instance measured thread.activity.append processing for that thread at:

    • average: 734 ms
    • minimum: 656 ms
    • maximum: 890 ms

    At the measured average, 526 events account for approximately 386 seconds of serial processing.

    The impact crosses threads because ProviderRuntimeIngestion feeds runtime events from every thread through one global sequential DrainableWorker. An unrelated turn’s turn.started and assistant-output events therefore wait behind the busy thread’s full-history projection scans. In the UI, the affected threads remained at “Connecting” despite their providers already working, and no assistant output appeared until the queue caught up.

    This also means disabling assistant streaming is only a partial workaround: tool-heavy activity can reproduce the same starvation. The practical workaround was to stop the high-activity session, wait for the queue to drain, and prefer fresh threads for tool-intensive work.

    The narrow fix appears to be event-specific shell-summary maintenance:

    • Do not call refreshThreadShellSummary for ordinary tool/context activity.
    • Do not call it for assistant messages.
    • Update latest_user_message_at directly for user messages.
    • Recompute pending-input state only for user-input lifecycle events.
    • Recompute approval and proposed-plan fields only from their relevant projections.
    • Reserve complete summary rebuilding for rare repair operations such as revert or reconciliation.

    Per-thread ingestion or lifecycle-event prioritization could provide defense in depth, but eliminating the hot-path full-history scans should address the primary backlog.

  3. dain commented on Aug 12, 2026

    @dain

    PR #5855 appears to fix the measured cause of this incident, though the globally sequential ingestion queue means other slow projection paths could still cause cross-thread head-of-line blocking.

  4. TimCrooker commented on Aug 21, 2026

    @TimCrooker
    Contributor

    Captured hard-freeze on current Nightly: 137.44 GB written by the T3 backend in 12.6 minutes

    This happened on:

    • T3 Code Nightly 0.0.34-nightly.20260821.1151
    • macOS 26.6.1 (25G76), Apple silicon, 48 GB RAM
    • Two active agent sessions; no Vitest, Turbo, or pnpm process during the final reproduction

    The machine became completely unresponsive and required a hard reset. Apple generated a disk-write diagnostic for the T3 backend immediately before the freeze:

    Event:            disk writes
    Writes:           137.44 GB of file-backed memory dirtied over 755 seconds
    Average:          182.12 MB/s
    T3 max footprint: 1283.53 MB
    ThermalPressure:  0
    

    The heaviest stack in that diagnostic was:

    node::sqlite::StatementSync::All              9,622 / 13,008 samples
    node::sqlite::StatementExecutionHelper::All   9,622 / 13,008 samples
    pwrite                                          5,237 / 13,008 samples
    

    The installed source, server trace, query plan, and live process samples all point to the same path:

    orchestration.command.thread.activity.append
      -> runProjectorForEvent
      -> applyThreadsProjection
      -> refreshThreadShellSummary
      -> ProjectionThreadActivityRepository.listByThreadId
    

    listByThreadId synchronously selects and decodes every activity row and payload_json for the thread. SQLite reports:

    SEARCH projection_thread_activities USING INDEX idx_projection_thread_activities_thread_sequence
    USE TEMP B-TREE FOR ORDER BY
    

    Live lsof monitoring showed the T3 backend repeatedly creating and deleting SQLite etilqs_* sort files, typically 37–63 MB each. One active thread had about 38,000 activity rows. It received 125 small activity events in five minutes, causing roughly 4.8 million historical-row visits and decodes for 0.122 MiB of new payload.

    The final failure timeline:

    • T3 backend repeatedly reached 100–142% CPU.
    • Disk traffic repeatedly reached 100–192 MiB/s.
    • Spotlight and endpoint scanners then reacted to the write storm and consumed the remaining CPU/I/O.
    • CPU idle fell to 1.98%; WindowServer reached about 59%.
    • Memory remained 84–85% free, swap remained 0, and thermal pressure remained 0.
    • The four-second system watcher stopped making progress at 09:48:25.
    • The one-second SQLite-temp watcher completed one delayed record at 09:49:48, then stopped.
    • The desktop remained frozen until the hard reset at 10:14:47.
    • No kernel panic or memory-pressure termination occurred.

    The initiating workload is T3's synchronous full-history scan and SQLite temp sort on routine activity events. macOS indexing and endpoint scanners amplify the resulting filesystem storm, but it begins inside the T3 backend and reproduces without repository test runners.

    For whichever linked fix lands (#5855, #7356, or #7486), a regression test should use a mature thread receiving routine activity events and assert that:

    1. Query count is constant with respect to thread history.
    2. Shell-summary maintenance does not query the full activity timeline.
    3. No SQLite temp-sort files are created for routine activity events.
    4. Sustained T3 disk writes remain below the macOS diagnostic threshold.

    I retained the Apple .diag, T3 stack samples, aligned process/system metrics, server trace, and SQLite temp-file observations. I can provide sanitized copies if needed.

  5. TimCrooker commented on Aug 21, 2026

    @TimCrooker
    Contributor

    One important impact note: this is not a one-off. Since adding more context to my normal T3 Code workflow, the machine has been freezing hard enough that I have to force-reboot it 1–5 times per day. I initially suspected RAM pressure or repository test runners, but the monitoring above traced the initiating workload back to T3 Code's backend projection and SQLite path. This is causing repeated work interruption and potential data loss, not just a transient performance slowdown.

  6. TimCrooker commented on Aug 21, 2026

    @TimCrooker
    Contributor

    Quick update after more monitoring and working with IT: we turned ThreatLocker off for an isolation test. The T3 evidence above still stands for the specific incident I captured — T3's projection/SQLite path wrote 137.44 GB and immediately preceded that freeze.

    I do want to narrow my earlier wording, though. T3 isn't the root cause of every hard reset. We later caught a separate memory storm while T3 wasn't running, with ThreatLocker growing to nearly 16 GB RSS, swap reaching about 29 GB, and disk traffic over 1 GB/s. So this looks like at least two different failure paths. The 1–5 forced reboots per day is the overall impact, not something I can attribute entirely to T3.

    One other correction: the earlier "swap stayed at 0" line came from a bad watcher parser, so swap for that reproduction should be treated as unknown.

  7. t3dotgg commented on Aug 22, 2026

    @t3dotgg
    Member

    Note

    🤖 GPT-5.6 Sol in Codex responding on behalf of Theo

    Closing as a duplicate of #4008. Both rescan the complete thread projection for each streamed text delta.

    Track the remaining work in #4008.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions