Skip to content

Prevent duplicate pending worker starts with invocation IDs and acknowledgements #696

Description

@jumski

Problem and scope

pgflow.ensure_workers() can start multiple live instances of the same Supabase Edge Function when startup takes longer than its debounce. Fix the observed startup race with a pending invocation ID and explicit worker acknowledgement, rather than a longer debounce.

This is the immediate 80/20 fix. Strict lifetime ownership, stale-worker fencing, and takeover after heartbeat loss are separate work in #697. This issue can ship independently; it does not promise that two runtime instances can never coexist.

Production evidence

Read-only inspection of cron history, HTTP responses, registrations, and workers on 2026-10-06, UTC, with pgflow 0.17.2:

Time Observation
17:11:08 pgflow_ensure_workers returned 10 rows: first startup batch.
17:11:18 The same job returned 10 rows: second batch before registration.
17:11:21–24 Twenty worker instances registered: two per function.
17:11:28 Another two functions received startup requests and subsequently had a third instance.
17:12:58 All ten functions had fresh 0.17.2 workers; most duplicates also passed a six-second heartbeat filter.

Cron ran every 5 seconds; each registration had 6-second debounce and start_mode = 'http'. Worker options: maxPollSeconds: 2, maxPgConnections: 3, visibilityTimeout: 5.

The twenty startup responses had HTTP 200, status: "started", and distinct worker IDs matching database rows. The first batch took roughly 13–16 seconds from invocation to registration. Evidence does not isolate cold boot, delivery, connection acquisition, or flow verification as the source of delay.

These were real concurrent instances, not merely stale rows in a 30-second query. Overlapping instances existed before the upgrade, so this is not established as a new 0.17.2 regression. No duplicate application side effects were demonstrated.

Cause

last_invoked_at records when cron sent a request, not whether the request is still pending. Once debounce expires, cron can send another request before the first worker registers.

FlowWorkerLifecycle.acknowledgeStart() verifies the flow before inserting its worker row. SupabasePlatformAdapter.ensureWorkerStarted() serializes startup within one runtime, but independent runtime instances do not share that guard.

Past worker history cannot tell whether a queued request will start a fresh runtime or reach an existing one. The handler must acknowledge the specific request.

Minimal protocol

  1. Reserve before sending. Under per-function transactional coordination, check healthy workers and pending invocation state. When startup is needed, record a unique invocation_id and database-clock expiry, and enqueue HTTP with that ID in the same transaction. Concurrent cron calls cannot create two current pending invocations. Preserve debounce as a retry throttle, not as the correctness mechanism.
  2. Reconcile every valid request. Authenticate as today and validate the invocation ID against the addressed function and current database state. Do this in the HTTP handler, including the first request that starts a worker, not module initialization. An ID is correlation, not authentication. A request can reach an already-running or surviving runtime.
  3. Accept readiness or dismiss failure. If an existing healthy worker handles the request, acknowledge it against that worker without starting another. Otherwise, finish initialization, then atomically validate/consume the still-current invocation and register the ready worker before entering its processing loop. Only one runtime can accept a token for a new worker. Receipt alone must not clear pending state; cron must never see a gap with neither pending invocation nor ready worker. Failure can dismiss the current invocation with a reason so cron can retry.
  4. Recover boundedly and idempotently. Failed delivery, hung startup, and lost acknowledgement must not block startup forever. Duplicate delivery/acknowledgement must not start another loop or clear a newer invocation. Revalidate at final readiness: expiry during initialization or a superseded ID rejects that start. Failed/expired/dismissed tokens cannot authorize fresh processing. If acknowledgement committed but its response was lost, consult existing worker/state rather than start again.

“Completed invocation” means the startup handshake completed, not that the worker stopped. Cron sends another request only when there is neither a healthy worker nor an unexpired pending invocation; completing a handshake is not by itself a reason to send another request.

Prefer a few fields on the existing registration plus small SQL operations. A separate invocation-history table, persisted received state, or general scheduler is not required. Choose a bounded pending timeout and explain the recovery trade-off; timing guesses alone must not replace token validation.

Boundaries and compatibility

  • pgflow cannot kill or evict a hosted Supabase runtime on demand. A dismissed request does not force Supabase to select a different runtime next time. Surviving runtimes must handle later valid requests correctly without needing redeployment.
  • Preserve separate function names, including intentional -b replicas on the same queue. Preserve enabled/deprecated behavior and keep process workers outside HTTP invocation scheduling.
  • Specify authenticated direct/manual invocation and local-development behavior. Any legacy tokenless path must not silently bypass the handshake while claiming its protection. Define SQL/worker rollout order and mixed-version limitations.
  • Do not add lifetime owner leases, generation fencing on task operations, or stricter drain semantics here. A previously ready worker can become stale, trigger replacement, and later resume; Design strict worker ownership and fencing for safe automatic replacement #697 owns that stronger problem. External side effects remain subject to existing retry/idempotency requirements.

Acceptance criteria

  • A deterministic regression holds initialization beyond debounce but within the pending timeout; repeated and concurrent cron calls enqueue only one request while it is pending. The test fails against current behavior.
  • Duplicate delivery across independent runtime instances accepts only one new worker for an invocation. Existing healthy runtimes acknowledge without another loop; receipt does not clear pending state, and readiness registration is atomic with acceptance.
  • Tests cover failed delivery/startup, lost acknowledgement, expiry during initialization, late requests, duplicate acknowledgement, and superseded IDs. Old requests cannot clear new pending state, and recovery needs neither forced runtime termination nor redeployment.
  • Tests preserve disabled registrations, distinct replicas, auth checks, and process-worker scheduling. Document manual/local behavior and safe mixed-version rollout.
  • Expose enough state to distinguish pending, ready, and retryable startup. Documentation states that this prevents redundant pending starts, not all overlap after heartbeat loss. Keep implementation smaller than the separate ownership/fencing proposal.

Source and related work

Mechanism checked against deployed SQL and published @pgflow/edge-worker@0.17.2; links track main, so inspect current source before implementation.

Split from the original broader scope: #697 tracks strict ownership and fencing during automatic replacement. It is a separate follow-up, not a blocker for this fix. Implementation order: #696 first; #697 starts only after #696 ships. #694 concerns an intermittent processing stall; no shared cause is established.

Activity

  1. changed the title [-]Enforce database-backed worker admission to prevent duplicate Supabase worker instances[/-] [+]Prevent duplicate pending worker starts with invocation IDs and acknowledgements[/+] on Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions