Repository navigation
Oversized Codex diff event exhausts desktop backend heap #10924
Description
Activity
Triage
Confirmed on current
main(6c583620ff7a, same commit as the nightly in the report). This is a real backend heap bug, not a duplicate of #8648 or #7075.What the code does
CodexAdapterbuilds every canonical event with the native payload attached, then mapsturn/diff/updatedontopayload.unifiedDiffas well:// apps/server/src/provider/Layers/CodexAdapter.ts (~985, ~1646) raw: { source, method, payload: event.payload ?? {} } // … type: "turn.diff.updated", payload: { unifiedDiff: payload.diff }
ProviderService.publishRuntimeEventwrites that object to the canonical provider log before publish (ProviderService.ts~919).EventNdjsonLogger.writethenJSON-encodes the whole event and keeps the resulting line (EventNdjsonLogger.ts~497, ~620).turn.diff.updatedis not in the transient skip set (that filter is for high-churn deltas). There is no per-record size cap. File rotation (10 MiB) and the flush watermark (1 MiB) run after serialization, so a 250 MiB CANON record is fully materialized first.The user-facing restart copy is a consequence, not a separate bug: after the Node process dies,
reconcileProviderSessionsmarks sessions whose provider process is gone as failed (serverRuntimeStartup.ts~339).Orchestration does not store
unifiedDiff.ProviderRuntimeIngestiononly usesturn.diff.updatedto dispatch a placeholder checkpoint (status: "missing",files: []). The large string exists for diagnostics (and an unused contract field), then gets serialized twice.A >250 MiB CANON record implies an underlying diff on the order of ~125 MiB written twice, plus the native log line, plus the
[ts] CANON: …copy. That is enough to push V8 over the edge on top of a live desktop heap.Related, not the same
- [Bug]: app-server stdin reader is quadratic in line length — OOM crash on large tool payloads #5389 (closed by fix(codex): avoid quadratic app-server input buffering #8605): quadratic app-server reader. Fixed. A comment there already saw large
turn/diff/updatedlines in provider logs; this issue is the later canonical duplicate + unbounded log serialize. - [Bug]: Desktop backend still SIGABRTs with V8 OOM after 9–18h on 0.0.35 (hydration bounds already present) #8648: idle / hydration OOM on a huge SQLite profile after hours. Different path.
- Codex thread resume can exhaust server heap on large histories #7075: Codex
thread/resumehistory. Different path.excludeTurns: trueis already on main (CodexSessionRuntime); fix(server): avoid loading Codex history on resume #8629 was closed as already landed.
No open PR bounds this logger path or removes the dual-field diff.
Suggested fix
- In
EventNdjsonLogger, omit or summarize oversized payloads beforeserializeEvent(keep method/type + byte count). That guards native and canonical across providers. - In
CodexAdapter, do not attach the fulldiffon bothraw.payloadandpayload.unifiedDiffforturn/diff/updated. - Add a regression test that a large
turn/diff/updatedcannot produce a multi-hundred-MiB CANON line.
Restarting the desktop app is only a temporary workaround.
- [Bug]: app-server stdin reader is quadratic in line length — OOM crash on large tool payloads #5389 (closed by fix(codex): avoid quadratic app-server input buffering #8605): quadratic app-server reader. Fixed. A comment there already saw large
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.acceptedfeature request acceptedfeature request acceptedvia-triageFiled through npx t3 triageFiled through npx t3 triage
on Sep 9, 2026 - added 8 commits that reference this issue
on Sep 9, 2026
What happened
The T3 Code desktop backend stops without warning. The UI then marks active work as failed and shows:
The fault happened three times over two days.
Diagnosis
The desktop backend runs out of V8 heap while it handles a large Codex
turn/diff/updatednotification.CodexAdapterplaces the same diff in two fields of the canonical event:raw.payload.diffpayload.unifiedDiffSource:
apps/server/src/provider/Layers/CodexAdapter.ts:985apps/server/src/provider/Layers/CodexAdapter.ts:1646ProviderServicesends this canonical event to the provider event logger before it publishes the event:apps/server/src/provider/Layers/ProviderService.ts:916EventNdjsonLoggerserializes the complete event and holds the resulting line in memory:apps/server/src/provider/Layers/EventNdjsonLogger.ts:497apps/server/src/provider/Layers/EventNdjsonLogger.ts:620In the observed crash, one redacted canonical log record exceeded 250 MiB. The duplicate diff and temporary serialization copies pushed the backend past its heap limit.
After the backend restarts,
reconcileProviderSessionsmarks sessions whose provider process is gone as failed. This creates the user-facing restart message:apps/server/src/serverRuntimeStartup.ts:338Expected behavior: T3 Code should bound, omit, or summarize large raw diff data before diagnostic serialization. One provider event should not stop the backend or unrelated work.
Steps to reproduce
Proposed source-level reproduction:
turn/diff/updatednotification with a large generated string inparams.diff.CodexAdapterconvert it to a canonical event.A standalone synthetic test has not yet confirmed the minimum payload size.
Version
0.0.41-nightly.20260909.1439, commit6c583620ff7aThis was the newest nightly, and its commit matched
mainwhen checked.Environment
Evidence
Sanitized provider log measurement:
No secrets, source paths, thread identifiers, thread titles, commands, query data, or payload content are included.
Related issues
Fix applied or workaround
Restarting the desktop app restores the backend for a time. Avoiding very large single-event diffs reduces the risk. No local data was changed.
Filed by
Codex, GPT-5, via
t3 triage