Skip to content

Design strict worker ownership and fencing for safe automatic replacement #697

Description

@jumski

Goal

Design strict database-authorized worker ownership across startup, failure, and automatic replacement: at most one current authorized polling owner per registered function. This is the stronger follow-up split from #696, which addresses redundant startup requests with an invocation/acknowledgement handshake.

Do not block #696 on this design. Ship the startup handshake independently. This issue is not a prerequisite for fixing the observed slow-start duplicate batch.

Why this is separate

A pending invocation prevents cron from repeatedly requesting startup while the previous request remains unresolved. It does not prove a previously running worker died when its heartbeat becomes stale.

Example: worker A pauses or loses database connectivity; its heartbeat expires; replacement B starts; A resumes. Preventing A from claiming new work or performing stale protected database writes needs ownership enforcement beyond startup acknowledgement.

No stale-owner side effects were demonstrated in the production incident documented in #696. This issue defines the stronger requested guarantee, not another confirmed incident or a new 0.17.2 regression.

Guarantee and platform limits

  • Guarantee at most one database-authorized polling owner per function, not exactly one physically live runtime at every instant. Zero owners during startup/recovery is allowed. Automatic recovery cannot distinguish a dead process from a paused or partitioned process with certainty.
  • pgflow cannot kill or evict a hosted Supabase runtime on demand. A surviving runtime can receive later HTTP requests; rejection must not assume Supabase routes the next request elsewhere. The operator identifies a new function deployment as the available way to move requests off an old version, not as a routine worker-count control mechanism or proof that old work stopped.
  • Separate registered function names, including intentional -b replicas serving the same queue, remain independent. Do not impose queue-wide exclusion or introduce a general autoscaler.
  • A lease alone does not fence stale work. External API calls already in flight cannot be revoked by PostgreSQL; this does not provide exactly-once external side effects. Existing retry/idempotency safeguards remain necessary.

Design work

  1. Atomic ownership and renewal. Prefer coordination through the existing registration where practical. Define owner ID, database-clock lease, and generation/fencing token. Startup and takeover compete atomically; renewal/release must match the current owner and generation. Integrate with Prevent duplicate pending worker starts with invocation IDs and acknowledgements #696's invocation IDs without treating invocation completion as lifetime ownership.
  2. Enforce ownership at the work boundary. Check ownership atomically with admission/claim, not merely in a preceding heartbeat. Audit task completion/failure and raw queue operations for stale writes and recovery consequences. A resumed old owner cannot renew, clear its successor, claim work, or mutate protected state as current owner.
  3. Define drain and takeover semantics. Current Supabase shutdown marks stopped concurrently with drain. Distinguish no-new-claims, draining, released, and physically terminated. Specify whether in-flight handlers may overlap a replacement and how their database results are treated; do not silently discard valid results or imply that abort cancels an external effect.
  4. Preserve compatibility and runtime behavior. Cover surviving isolates receiving fresh requests, disabled/deprecated workers, intentional replicas, and process workers. Define migration/rollout handling for older binaries without fencing; do not claim strict exclusion during an unsafe mixed-version deployment.

Prefer the smallest protocol that proves the chosen invariant. Session-level advisory locks alone are unsuitable for transaction-pooled connections and do not fence stale external effects. No arbitrary replica-count API or separate scheduling service is requested.

Acceptance criteria

  • State the safety and availability contract, including claim-time ownership, permitted drain overlap, recovery gaps, and external side-effect limits.
  • Deterministic tests pause owner A, expire its lease, admit B, then resume A; A cannot claim or perform protected stale writes, renew, or release B's ownership.
  • Concurrent admission/takeover, lost renewal responses, hard death, graceful drain, deprecation, and disabled registrations recover without permanent starvation or forced runtime termination.
  • Two distinct functions serving one queue retain intentional parallelism. Process-worker behavior and mixed-version rollout are specified and tested.
  • Monitoring distinguishes current authorized owner from draining, stale, and historical runtimes. Implementation preserves task recovery and documents what fencing cannot guarantee.

Code and related work

Findings derive from the published @pgflow/edge-worker@0.17.2 source; inspect current main before implementation.

Immediate fix: #696 — pending startup invocation and acknowledgement. This issue: stronger ownership during failure/replacement. Implementation order: #696 first; this issue starts only after #696 ships, and may reuse its invocation state. #694 concerns a separate processing stall; no shared cause is established.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions