Skip to content

V8 OOM crash → white screen after extended sessions (Linux) #1686

Description

@mihai2mn

Description

T3 Code (Alpha) desktop app on Linux freezes and goes to a white screen during extended sessions. The window frame stays alive but the renderer is dead.

Root Cause (from logs)

The Electron renderer process hits V8's heap limit (~3.7 GB) and crashes with OOM:

[585758:0x1144004e0000] Mark-Compact (reduce) 3746.0 (4003.8) -> 3746.0 (4002.8) MB, pooled: 0.0 MB, 540.68 / 0.00 ms (average mu = 0.058, current mu = 0.000) last resort; GC in old space requested
V8 javascript OOM (Ineffective mark-compacts near heap limit).
Process 585758 (t3-code-desktop) dumped core.
Consumed 3h 22min 35.026s CPU time over 1h 6min 45.825s wall clock time, 22.3G memory peak, 448.4M memory swap peak.

V8 GC runs "last resort" compaction twice but reclaims 0 bytes — the entire 3.7 GB heap is live/reachable.

Reproduction

  1. Open T3 Code desktop on Linux
  2. Run an extended session (~1 hour) with heavy tool use (many Bash calls, large file reads, code diffs)
  3. App freezes, then goes to white screen

Happens intermittently but reproducibly on long sessions with large tool outputs.

Environment

  • App: T3 Code (Alpha), installed from AUR (t3code-bin) at /opt/t3code-bin/
  • OS: CachyOS (Arch-based), Linux 6.19.10-1-cachyos, x86-64
  • Memory: System has sufficient RAM; the issue is V8 heap exhaustion in the renderer process

Expected Behavior

The app should either:

  • Evict old conversation turns from the renderer's DOM/React state to stay within heap limits
  • Gracefully recover (reload the renderer) instead of showing a dead white screen
  • Show an error message suggesting the user start a fresh conversation

Workaround

Using the CLI (claude) instead of the desktop app avoids the issue entirely since there's no Electron layer.

Log Locations

  • Desktop main log: ~/.t3/userdata/logs/desktop-main.log
  • Server child log: ~/.t3/userdata/logs/server-child.log
  • Provider logs: ~/.t3/userdata/logs/provider/_global.log
  • Core dump: captured by systemd-coredump

Activity

  1. Zimtente commented on Apr 3, 2026

    @Zimtente

    experiencing the same. the performance also got quite bad with recent release.

    Popos 22.04 LTS

  2. keir-whitehead commented on Apr 5, 2026

    @keir-whitehead

    +1 here

    52 renderer crashes in 7 days, all with identical signatures:

    RAM 128gb
    Crash frequency 6-19 per day
    Root Cause: V8 Garbage Collection Crash
    The crash stacks consistently show the Chromium renderer dying during V8 garbage collection

  3. moon-strider commented on Apr 5, 2026

    @moon-strider

    Same thing on newest MacOS (M1 Pro, 16GB RAM) — relaunch fixes it, appears seemingly every hour or two. I couldn't pinpoint an exact cause, but it could have something to do with bringing the app in and out of focus, because it never happened while the app was in focus

  4. added a commit that references this issue on Apr 8, 2026
    4457096
  5. benhook1013 commented on Jul 6, 2026

    @benhook1013

    I had to throw away a long lived thread to avoid crashes, was not an issue before recent changes, assume the tracking of entire thread's history including the sidebar to jump to position is responsible.

  6. nassimna commented on Jul 22, 2026

    @nassimna
    Contributor

    Additional reproduction on 0.0.29-nightly.20260722.878

    This is still reproducible on the July 22 nightly. In this case the desktop window opens, but input becomes effectively unresponsive while the renderer repeatedly performs last-resort GC, then the renderer crashes. I observed the same failure twice in one session.

    Crash evidence

    Mark-Compact (reduce) 3479.6 (...) MB ... last resort; GC in old space requested
    V8 javascript OOM (CALL_AND_RETRY_LAST)
    

    The renderer core dump was SIGTRAP with:

    VmRSS: 3950108 kB
    VmHWM: 4104272 kB
    Threads: 24
    

    A previous renderer crash in the same session peaked at about 4.01 GB RSS, with the same signature.

    Profile size at reproduction

    • ~/.t3/userdata/state.sqlite: 187 MB
    • projection_thread_activities: 20,939 rows, about 63.46 MiB of payload_json + summary
    • orchestration_events: 27,265 rows
    • Largest single activity stream: 3,688 rows / 23.78 MiB
    • Largest individual activity payload observed: about 1.01 MiB
    • Projection cursors were fully caught up to the latest event sequence, and PRAGMA quick_check returned ok

    This profile is much smaller than the multi-gigabyte backend-OOM reproductions discussed in #996, yet its renderer still expands to roughly 4 GB and crashes.

    Suspected path

    The packaged server code includes an activity snapshot query equivalent to the query in apps/server/src/orchestration/Layers/ProjectionSnapshotQuery.ts that selects all rows from projection_thread_activities, ordered, without a WHERE window or LIMIT. There is also an unbounded per-thread activity query.

    The on-disk activity JSON expanding from about 63 MiB to a greater-than-3.4-GB live V8 heap strongly suggests eager activity hydration/retention is involved. That relationship is an inference from the query and measurements, not a heap-object trace.

    This looks related to #996 and the pagination/lazy-history work in #3510. The useful distinction here is that the crashing process is the Electron renderer, not the backend child, and the user-visible symptom starts as a window that cannot be clicked before the renderer dies.

    Environment

    • T3 Code Nightly 0.0.29-nightly.20260722.878
    • Arch Linux x86_64
    • X11 / bspwm
    • Official AppImage payload installed through the AUR nightly package
    • Electron launched with --no-sandbox --ozone-platform-hint=auto

    Restarting does not resolve it because the retained history is loaded again. I have not deleted or compacted the profile, so there is no verified non-destructive workaround from this reproduction.

  7. benhook1013 commented on Jul 23, 2026

    @benhook1013

    I found a context-preserving workaround for what appears to be this same eager activity-hydration path.

    This does not fix the renderer bug, and it deliberately removes old UI scrollback, but it restored performance for me without losing provider/model continuation. My affected thread's projected UI history was reduced from 278 MB to 50.86 MB while retaining the latest 12 complete user prompts.

    I deleted rows only from:

    • projection_thread_messages
    • projection_thread_activities

    I left these untouched:

    • orchestration_events
    • projection_turns
    • projection_thread_sessions
    • provider_session_runtime and its resume cursor
    • projection cursors/state

    Afterward, the event count, turn count, and provider resume cursor matched the pre-trim backup exactly, and PRAGMA integrity_check returned ok. Old visual scrollback disappeared, but the continuing model session retained its context. A full projection rebuild may recreate the old rows from retained events.

    If you have an external CLI agent available, you could give it this task:

    Fully stop T3 Code and inspect ~/.t3/userdata/state.sqlite. Do not modify anything until you have identified the exact thread and confirmed the current schema. Create a consistent backup using SQLite's .backup command—not cp—and verify it with PRAGMA quick_check.

    Measure projection_thread_messages and projection_thread_activities by thread. For the affected thread, rank role='user' messages by created_at DESC, message_id DESC. Evaluate retained payload size at each prompt boundary and select a reasonable target.

    In one guarded transaction, delete only message and activity projection rows whose created_at is older than the chosen user-prompt timestamp. This must retain that prompt and every message/activity after it, so long-running turns remain whole.

    Do not delete or alter orchestration events, turns, thread sessions, provider runtime/session records, resume cursors, or projection-state watermarks. Afterward, verify the retained prompt count and payload size, run PRAGMA integrity_check, and compare event counts, turn counts, and the provider resume cursor with the backup. Keep the backup and restart T3 Code. Do not run VACUUM while T3 is running.

    Since your entire activity projection is about 63 MiB rather than my 278 MB single-thread projection, this isn't proof it will resolve your crash—but it is a reversible way to test the suspected hydration path while preserving the actual continuation state.

  8. t3-code commented on Aug 20, 2026

    @t3-code
    Contributor

    closing as completed. #5148 and #5147 bounded renderer/server hydration and added recovery from renderer OOM. Please open a fresh report with current-build diagnostics if this recurs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions