Skip to content

[Bug]: V2 storage cleanup never frees worktrees for completed/cancelled/interrupted/rolled_back threads #15146

Description

@ANSHSINGH050404

What happened

V2 threads that finish as completed/cancelled/interrupted/rolled_back keep their worktrees forever. storageCleanupThreadIdle only returns true for status idle or failed, so any thread whose shell status mirrors a terminal run status other than failed is never eligible for worktree cleanup. Disk usage grows unboundedly, especially on remote servers.

Diagnosis

Grounded in source at upstream/main@e8545b293b:

  • packages/contracts/src/orchestrationV2.ts:445-456 OrchestrationV2RunStatus includes preparing, queued, starting, running, waiting, completed, interrupted, failed, cancelled, rolled_back.
  • packages/contracts/src/orchestrationV2.ts:1661-1664 OrchestrationV2ShellThreadStatus is idle plus RunStatus, so thread.status can be completed/interrupted/cancelled/rolled_back.
  • apps/server/src/orchestration-v2/ProjectionStore.ts:1485-1499 shellStatusFromStoredRunStatus mirrors latest_run_status verbatim, including completed/interrupted/cancelled/rolled_back.
  • apps/server/src/storageCleanup.ts:83-92 storageCleanupThreadIdle requires thread.activeRunId === null AND (status === idle OR status === failed). Terminal threads with status completed/cancelled/interrupted/rolled_back and activeRunId null fail the predicate and are skipped at :220 and :317.

V1 allowed cleanup whenever session was null/stopped and latestTurn was not running, regardless of terminal state (024d495:apps/server/src/storageCleanup.ts:82-91). The idle-plus-failed gate is introduced by the orchestrator V2 merge de34391.

Steps to reproduce

Logic repro, no provider CLI needed:

  1. Construct an OrchestrationV2ThreadShell with branch and worktreePath set, activeRunId null, status completed, pendingBackgroundTasks [], pendingRuntimeRequest null, no queued turn.
  2. Call storageCleanupThreadIdle(thread, Date.now()).
  3. Result is false, so the worktree candidate is filtered out. Same for cancelled, interrupted, rolled_back. Only idle and failed return true.

Observable product repro:

  1. Run a V2 thread to completion (status completed, no active run).
  2. Wait past worktree retention/inactivity threshold.
  3. Worktree directory remains; storage cleanup logs skip it as non-idle.

Version

upstream/main@e8545b293b, post orchestrator V2 merge de34391; checked 2026-10-03.

Environment

Windows x64 repo checkout, verified via git show upstream/main. Server-side defect, applies on any OS/host, especially remote servers with many finished threads.

Evidence

  • apps/server/src/storageCleanup.ts:83-92 predicate, :220 and :317 call sites
  • packages/contracts/src/orchestrationV2.ts:445-456 RunStatus, :1661-1664 ShellThreadStatus
  • apps/server/src/orchestration-v2/ProjectionStore.ts:1485-1499 status mirroring
  • V1 baseline 024d495:apps/server/src/storageCleanup.ts:82-91

Focused verification via git show upstream/main only; no repo-wide lint/typecheck per repo rules. Targeted vpr test run failed on runner config in this checkout.

Related issues

Fix applied or workaround

None. Possible direction: treat terminal statuses completed/cancelled/interrupted/rolled_back as idle-eligible once activeRunId is null and background/queued guards pass, or add explicit terminal-state handling with inactivity clock. Not tested.

Filed by

opencode (muse-spark-1.3-contributor-free) via audit session

Activity

  1. juliusmarminge commented on Oct 3, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Triage

    Thanks @ANSHSINGH050404 for the precise, source-grounded report! This is a real regression in the V2 worktree-cleanup idle check. It's not a duplicate of #14742, #14843, #13836, or #14966 (all open, and each about a different cleanup path).

    What happens

    OrchestrationV2ThreadShell.status holds the latest run status. shellStatusFromStoredRunStatus passes completed, interrupted, failed, cancelled, and rolled_back straight through, and the shell snapshot publishes that value as status. active_run_id is only set for preparing, starting, and running, so a finished thread has activeRunId === null and a terminal status. idle only shows up when a thread has never run (latest_run_status is null).

    storageCleanupThreadIdle in apps/server/src/storageCleanup.ts still requires status === "idle" || status === "failed". Both places that call it (the candidate filter and the recheck just before deletion) therefore skip completed, cancelled, interrupted, and rolled_back threads before worktreeAfterDays, worktreeOnMerge, or worktreeUnchanged can apply. worktreeOnDelete doesn't use this check, so cleanup when a thread is deleted still works. The other three rules are off by default, so this only shows up where they've been turned on.

    V1 cleaned a worktree whenever the session was null or stopped and the latest turn wasn't running. The orchestrator V2 merge (#2829, de34391427) replaced that with the idle/failed gate. The current tests only cover active statuses (running, starting, preparing, waiting, queued). They don't require the other terminal statuses to stay ineligible. queued and waiting still need to stay excluded, because active_run_id doesn't cover them.

    One small correction to the repro: the skip is a silent continue, so no "non-idle" line is logged. storage cleanup skipped worktree only appears if a later step fails.

    Live work is still protected after this check. Any provider session that isn't stopped and has its cwd inside the worktree blocks removal, and idle sessions are released after 30 minutes. Dirty trees, ignored files other than node_modules, and live terminals also still block removal.

    Likely fix area

    Once activeRunId is null and the existing background-task, runtime-request, and queued-turn guards pass, completed, interrupted, cancelled, and rolled_back could be treated the same as failed, with cases for those statuses added to apps/server/src/storageCleanup.test.ts. A maintainer will decide on the fix direction.

  2. added
    bugSomething is broken or behaving incorrectly.
    via-triageFiled through npx t3 triage
    on Oct 3, 2026
  3. kamil-koryciorz commented on Oct 6, 2026

    @kamil-koryciorz

    Reproduced on a packaged nightly, not only in source.

    Version: T3 Code Nightly 0.0.46-nightly.20261005.2689 (desktop), macOS 26.7.1 (Apple Silicon)

    Setup: Settings → Storage with "Delete merged worktrees", "Delete unchanged worktrees" and "Delete inactive worktrees" (14 days) on, no project override. Test project: a small GitHub repo with main at the same commit as origin/main.

    Steps

    1. Launch a worktree thread (t3_thread_launch, Claude) that runs one turn changing nothing and completes. Thread status reads completed, activeRunId null, no pending requests.
    2. Wait 30 min so the provider session reaches stopped (checked in orchestration_v2_projection_provider_sessions).
    3. Flip any cleanup setting to trigger a sweep.

    Actual: worktree stays. Trace shows StorageCleanup.sweep → cleanWorktrees running, but no StorageCleanup.ignoredFiles span for that worktree, i.e. it is filtered before the git checks.

    Control: a thread launched with no message (status idle, never run) in the same project lost its clean worktree on the very next sweep. Same settings, same repo, same minute.

    So on this build the three rules only ever act on never-run or failed threads, as the issue describes. #15150 would cover it.

    Written with Claude Fable 5.1 through T3's MCP tools; the human owner reviewed before posting.

  4. Info-Cado commented on Oct 6, 2026

    @Info-Cado

    Reproduced at scale on Linux. Still present on main at 9bd1d8009a (apps/server/src/storageCleanup.ts:89).

    Version: T3 Code Nightly 0.0.46-nightly.20261005.2667 (Linux AppImage)

    Settings: storageCleanup: { worktreeAfterDays: 5, worktreeOnMerge: true, worktreeOnDelete: true, worktreeUnchanged: true }, no project overrides.

    One private project, 28 linked worktrees under ~/.t3/worktrees/<project>/, grouped by the status of the top-level thread that owns each one:

    Thread status Worktrees PR merged
    completed 23 22
    idle (never ran in V2, createdBy: "system") 3 3
    running 2 0

    The 23 completed threads all have activeRunId: null, no pending runtime request and no pending background tasks. 22 of them have a PR with pullRequests[].snapshot.state: "merged" and were settled between 2026-10-03 and 2026-10-05.

    Traces (server.trace.ndjson*, 2026-10-06 12:47–13:22 UTC): 32 StorageCleanup.sweep traces. Every sweep that reached the Git checks ran StorageCleanup.containsProjectRoot and StorageCleanup.ignoredFiles for the same 9 worktrees across all projects (by git.cwd). All 9 belong to never-run idle threads. No worktree of a completed thread appears in any sweep.

    What #15150 would not change here. So nobody expects these worktrees to go away once it merges: on this machine, other checks still keep every one of them.

    Also, #15080 moves this predicate to worktreeThreadState.ts and keeps the idle/failed gate, so it would carry this bug if it lands before #15150.

    Evidence was collected read-only: settings.json, sqlite3 -readonly on statev2.sqlite, the server traces, and local git reads in each worktree.

    Investigated and written by Claude Opus 5.5 through T3 Code's Claude Code harness, on behalf of the reporter.

  5. SpyrosPsarras commented on Oct 6, 2026

    @SpyrosPsarras

    I see a side effect of this on v0.0.46-nightly.20261003.2632. Because worktrees are never freed, they pile up: one project has 56 under ~/.t3/worktrees. The cleanup sweep then gets expensive and runs almost constantly.

    From one hour of the server trace:

    • StorageCleanup.cleanWorktrees ran 38 times, back to back. Each run took 28–186 s, 61 s on average.
    • The trace logged 34.7k runGitCommand spans in total. Single git calls took up to 16 s, and some hit Git command timed out in GitVcsDriver.resolveRepositoryPaths.
    • VCS status polling for the same worktrees adds load: VcsStatusBroadcaster.refreshRemoteStatus took up to 22.9 s, and branchPullRequest took up to 17.2 s.

    The sweep removes nothing, because of the eligibility check described above. It still walks every worktree each time. On this server that load makes the SQLite stalls in #14701 worse. Throttling the sweep, or skipping worktrees whose thread status cannot change eligibility, would remove most of this git traffic even before the eligibility fix lands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions