Skip to content

fix(workspaces): recover from watcher startup failures - #2995

Closed
EivindKjosbakken wants to merge 1 commit into
generalaction:mainfrom
EivindKjosbakken:fix/workspace-registry-fsevents-recovery
Closed

EivindKjosbakken wants to merge 1 commit into
generalaction:mainfrom
EivindKjosbakken:fix/workspace-registry-fsevents-recovery

Conversation

@EivindKjosbakken

Copy link
Copy Markdown

Description

Recover the workspace registry when a native filesystem watcher cannot start.

Previously, a rejected WatchHandle.ready() promise escaped as an unhandled rejection and terminated the workspace-registry worker. After the generic supervision budget was exhausted, new workspace activation stayed unavailable until the entire desktop app was restarted.

This change:

  • observes watcher startup failures and keeps the polling freshness floor active;
  • retries failed native watches with bounded exponential backoff;
  • supervises the workspace-registry worker indefinitely with a 30-second maximum delay;
  • clears all retry state when targets disappear or the scheduler is disposed.

Related issues

No existing issue found.

Testing

  • pnpm --filter @emdash/core run format
  • pnpm --filter @emdash/core run typecheck
  • pnpm --filter @emdash/core run lint
  • pnpm --filter @emdash/core exec vitest run --maxWorkers=1 (203 files, 1,369 tests)
  • pnpm run build (9 projects)
  • packaged an arm64 macOS Canary build
  • local Codex review: clean after addressing its polling-fallback finding

Screenshot/Recording (if applicable)

Not applicable; this is worker recovery behavior.

Checklist
  • I kept this PR small and focused
  • I ran a self-review before opening this PR
  • I ran the relevant local checks or explained why not
  • I updated docs when behavior or setup changed, or docs were not required
  • I added or updated tests when behavior changed, or explained why not
  • I only added comments where the logic is not obvious
  • I used Conventional Commits for the commit message and PR title

@greptile-apps

greptile-apps Bot commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

The PR makes workspace-registry recovery resilient to native watcher startup failures and exhausted worker restart budgets.

  • Observes watcher readiness failures while retaining periodic polling.
  • Retries native watcher creation with capped exponential backoff.
  • Clears watcher retry timers and attempt state when targets disappear or the scheduler is disposed.
  • Gives the workspace-registry worker indefinite failure supervision with a 30-second maximum delay.
  • Adds focused tests for polling fallback, watcher reattachment, and worker supervision behavior.

Confidence Score: 5/5

The PR appears safe to merge, with watcher recovery, polling fallback, retry cleanup, and worker shutdown behavior remaining coherent.

Failed watcher creations are released and retried through fresh resource-cache entries while polling continues independently, and indefinite worker supervision remains cancellable during shutdown.

Important Files Changed

Filename Overview
packages/core/src/runtimes/workspace-registry/node/scan/scheduler.ts Adds guarded watcher readiness handling, bounded retry state, and complete timer cleanup without disrupting the independent polling floor.
packages/core/src/runtimes/workspace-registry/node/scan/scheduler.test.ts Adds coverage for polling during watcher failure and event delivery after a failed watch is successfully reattached.
packages/core/src/runtimes/workspace-registry/node/worker-spec.ts Replaces bounded default worker supervision with intentional indefinite on-failure retries capped at 30 seconds.
packages/core/src/runtimes/workspace-registry/node/worker-spec.test.ts Verifies the worker鈥檚 restart policy, initial delay, and indefinitely repeated maximum delay.

Sequence Diagram

sequenceDiagram
  participant S as Scan Scheduler
  participant W as Watch Service
  participant P as Polling Floor
  S->>W: watch(target)
  W-->>S: ready() rejects
  S->>W: release failed handle
  S->>S: schedule bounded backoff
  P->>S: request freshness scan
  S->>W: retry watch(target)
  W-->>S: ready() resolves
  W-->>S: filesystem events resume
Loading

Reviews (1): Last reviewed commit: "fix(workspaces): recover from watcher st..." | Re-trigger Greptile

@jschwxrz

jschwxrz commented Sep 7, 2026

Copy link
Copy Markdown
Collaborator

closing as resolved

@jschwxrz jschwxrz closed this Sep 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants