Steps to reproduce
- Run a self-hosted server (
t3 serve) with the Codex provider.
- Start a thread and have the agent use the browser preview to fetch and read
web pages, so tool.completed results contain full page content.
- Let it run for a while — in my case roughly an hour of ordinary browsing work
with a single agent.
I don't have a minimal deterministic repro, sorry. The trigger seems to be a
single JSON-RPC line on the app-server stdio pipe that is large (megabytes) and
therefore arrives across many chunks. Any tool result big enough to be delivered
in many reads should do it; page content is just the easiest way to get there.
Expected behavior
Reading a large message off the app-server pipe costs time proportional to the
size of the message.
Actual behavior
The server's CPU goes to ~100% of a core and heap climbs until Node aborts:
FATAL ERROR: Reached heap limit Allocation failed - JavaScript heap out of memory
1: 0xe42d60 node::OOMErrorHandler(char const*, v8::OOMDetails const&)
2: 0x121ddd0 v8::Utils::ReportOOMFailure(v8::internal::Isolate*, char const*, ...)
3: 0x121e0a7 v8::internal::V8::FatalProcessOutOfMemory(...)
5: 0x1465299 v8::internal::Heap::CollectGarbage(...)
6: 0x1439948 v8::internal::HeapAllocator::AllocateRawWithLightRetrySlowPath(...)
7: 0x143a875 v8::internal::HeapAllocator::AllocateRawWithRetryOrFailSlowPath(...)
and systemd records:
t3code.service: Main process exited, code=dumped, status=6/ABRT
t3code.service: Failed with result 'core-dump'.
t3code.service: 8.8G memory peak
The clients disconnect when the process dies. systemd restarted it and it
happened again on the same workload.
Where it comes from
packages/effect-codex-app-server/src/protocol.ts (v0.0.31, around line 354),
in the stdin reader:
Ref.modify(remainder, (current) => {
const combined = current + chunk;
const lines = combined.split("\n");
const nextRemainder = lines.pop() ?? "";
return [lines.map((line) => line.replace(/\r$/, "")), nextRemainder] as const;
})
Every chunk concatenates onto the whole pending buffer and re-splits all of it.
For one line arriving across N chunks that's O(n²) in both copying and
allocation. A ~50 MB line in 64 KB chunks works out to roughly 20 GB of copying
to receive 50 MB.
It only shows up when a single line is large. Normal coding tool results are
small enough that the accumulated buffer never gets big, which I assume is why
this hasn't come up before.
Evidence
V8 CPU profile of the process (Profiler.start, 25 s, sampled at 100 µs) while
reproducing, aggregated by self time:
96.8% (anonymous) bin.mjs:68926 -> protocol.ts:354 via the shipped sourcemap
2.5% (garbage collector)
0.0% everything else
process.memoryUsage() at the same moment:
rss 1955.7 MB
heapUsed 1743.3 MB
external 3.8 MB
arrayBuffers 0.2 MB
So it is JS heap, not buffers.
Possible fix
Holding the pending bytes as the chunks that produced them and joining only when
a newline actually arrives makes line assembly linear — a chunk with no newline
becomes an array push instead of a full buffer copy:
const remainder = yield* Ref.make<ReadonlyArray<string>>([]);
Ref.modify(remainder, (current) => {
if (!chunk.includes("\n")) {
return [[] as ReadonlyArray<string>, [...current, chunk]] as const;
}
const combined = current.length === 0 ? chunk : current.join("") + chunk;
const lines = combined.split("\n");
const nextRemainder = lines.pop() ?? "";
return [
lines.map((line) => line.replace(/\r$/, "")),
nextRemainder.length === 0 ? [] : [nextRemainder],
] as const;
})
The stream-end path also needs Effect.map((pending) => pending.join("")) before
handling the trailing partial line.
I've been running this on a local fork since the crash. Under the same browsing
workload the profile no longer shows this function at all, heap now sawtooths and
recovers instead of climbing, and the OOM hasn't recurred. Package typecheck and its 20 tests pass, and the
server suite (1771 passed / 7 skipped) is unaffected. Happy to open a PR if
that's useful, or equally happy for you to take the idea and do it your own way —
I know you're not looking for contributions right now.
Note
I can see #2829 is rewriting the orchestration layer, and I had a look to check
whether this was already covered. As far as I can tell this code is in the
app-server protocol package rather than the orchestration layer, so it may
outlive V2 — but you'd know far better than I would.
Workaround
None that I found short of patching. Avoiding large tool results (not using the
browser preview on content-heavy pages) keeps lines small enough to stay clear
of it.
Steps to reproduce
t3serve) with the Codex provider.web pages, so
tool.completedresults contain full page content.with a single agent.
I don't have a minimal deterministic repro, sorry. The trigger seems to be a
single JSON-RPC line on the app-server stdio pipe that is large (megabytes) and
therefore arrives across many chunks. Any tool result big enough to be delivered
in many reads should do it; page content is just the easiest way to get there.
Expected behavior
Reading a large message off the app-server pipe costs time proportional to the
size of the message.
Actual behavior
The server's CPU goes to ~100% of a core and heap climbs until Node aborts:
and systemd records:
The clients disconnect when the process dies. systemd restarted it and it
happened again on the same workload.
Where it comes from
packages/effect-codex-app-server/src/protocol.ts(v0.0.31, around line 354),in the stdin reader:
Every chunk concatenates onto the whole pending buffer and re-splits all of it.
For one line arriving across N chunks that's O(n²) in both copying and
allocation. A ~50 MB line in 64 KB chunks works out to roughly 20 GB of copying
to receive 50 MB.
It only shows up when a single line is large. Normal coding tool results are
small enough that the accumulated buffer never gets big, which I assume is why
this hasn't come up before.
Evidence
V8 CPU profile of the process (
Profiler.start, 25 s, sampled at 100 µs) whilereproducing, aggregated by self time:
process.memoryUsage()at the same moment:So it is JS heap, not buffers.
Possible fix
Holding the pending bytes as the chunks that produced them and joining only when
a newline actually arrives makes line assembly linear — a chunk with no newline
becomes an array push instead of a full buffer copy:
The stream-end path also needs
Effect.map((pending) => pending.join(""))beforehandling the trailing partial line.
I've been running this on a local fork since the crash. Under the same browsing
workload the profile no longer shows this function at all, heap now sawtooths and
recovers instead of climbing, and the OOM hasn't recurred. Package typecheck and its 20 tests pass, and the
server suite (1771 passed / 7 skipped) is unaffected. Happy to open a PR if
that's useful, or equally happy for you to take the idea and do it your own way —
I know you're not looking for contributions right now.
Note
I can see #2829 is rewriting the orchestration layer, and I had a look to check
whether this was already covered. As far as I can tell this code is in the
app-server protocol package rather than the orchestration layer, so it may
outlive V2 — but you'd know far better than I would.
Workaround
None that I found short of patching. Avoiding large tool results (not using the
browser preview on content-heavy pages) keeps lines small enough to stay clear
of it.