You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Design strict worker ownership and fencing for safe automatic replacement #697
Design strict database-authorized worker ownership across startup, failure, and automatic replacement: at most one current authorized polling owner per registered function. This is the stronger follow-up split from #696, which addresses redundant startup requests with an invocation/acknowledgement handshake.
Do not block #696 on this design. Ship the startup handshake independently. This issue is not a prerequisite for fixing the observed slow-start duplicate batch.
Why this is separate
A pending invocation prevents cron from repeatedly requesting startup while the previous request remains unresolved. It does not prove a previously running worker died when its heartbeat becomes stale.
Example: worker A pauses or loses database connectivity; its heartbeat expires; replacement B starts; A resumes. Preventing A from claiming new work or performing stale protected database writes needs ownership enforcement beyond startup acknowledgement.
No stale-owner side effects were demonstrated in the production incident documented in #696. This issue defines the stronger requested guarantee, not another confirmed incident or a new 0.17.2 regression.
Guarantee and platform limits
Guarantee at most one database-authorized polling owner per function, not exactly one physically live runtime at every instant. Zero owners during startup/recovery is allowed. Automatic recovery cannot distinguish a dead process from a paused or partitioned process with certainty.
pgflow cannot kill or evict a hosted Supabase runtime on demand. A surviving runtime can receive later HTTP requests; rejection must not assume Supabase routes the next request elsewhere. The operator identifies a new function deployment as the available way to move requests off an old version, not as a routine worker-count control mechanism or proof that old work stopped.
Separate registered function names, including intentional -b replicas serving the same queue, remain independent. Do not impose queue-wide exclusion or introduce a general autoscaler.
A lease alone does not fence stale work. External API calls already in flight cannot be revoked by PostgreSQL; this does not provide exactly-once external side effects. Existing retry/idempotency safeguards remain necessary.
Design work
Atomic ownership and renewal. Prefer coordination through the existing registration where practical. Define owner ID, database-clock lease, and generation/fencing token. Startup and takeover compete atomically; renewal/release must match the current owner and generation. Integrate with Prevent duplicate pending worker starts with invocation IDs and acknowledgements #696's invocation IDs without treating invocation completion as lifetime ownership.
Enforce ownership at the work boundary. Check ownership atomically with admission/claim, not merely in a preceding heartbeat. Audit task completion/failure and raw queue operations for stale writes and recovery consequences. A resumed old owner cannot renew, clear its successor, claim work, or mutate protected state as current owner.
Define drain and takeover semantics. Current Supabase shutdown marks stopped concurrently with drain. Distinguish no-new-claims, draining, released, and physically terminated. Specify whether in-flight handlers may overlap a replacement and how their database results are treated; do not silently discard valid results or imply that abort cancels an external effect.
Preserve compatibility and runtime behavior. Cover surviving isolates receiving fresh requests, disabled/deprecated workers, intentional replicas, and process workers. Define migration/rollout handling for older binaries without fencing; do not claim strict exclusion during an unsafe mixed-version deployment.
Prefer the smallest protocol that proves the chosen invariant. Session-level advisory locks alone are unsuitable for transaction-pooled connections and do not fence stale external effects. No arbitrary replica-count API or separate scheduling service is requested.
Acceptance criteria
State the safety and availability contract, including claim-time ownership, permitted drain overlap, recovery gaps, and external side-effect limits.
Deterministic tests pause owner A, expire its lease, admit B, then resume A; A cannot claim or perform protected stale writes, renew, or release B's ownership.
Concurrent admission/takeover, lost renewal responses, hard death, graceful drain, deprecation, and disabled registrations recover without permanent starvation or forced runtime termination.
Two distinct functions serving one queue retain intentional parallelism. Process-worker behavior and mixed-version rollout are specified and tested.
Monitoring distinguishes current authorized owner from draining, stale, and historical runtimes. Implementation preserves task recovery and documents what fencing cannot guarantee.
Code and related work
Findings derive from the published @pgflow/edge-worker@0.17.2 source; inspect current main before implementation.
Immediate fix:#696 — pending startup invocation and acknowledgement. This issue: stronger ownership during failure/replacement. Implementation order: #696 first; this issue starts only after #696 ships, and may reuse its invocation state. #694 concerns a separate processing stall; no shared cause is established.
Goal
Design strict database-authorized worker ownership across startup, failure, and automatic replacement: at most one current authorized polling owner per registered function. This is the stronger follow-up split from #696, which addresses redundant startup requests with an invocation/acknowledgement handshake.
Do not block #696 on this design. Ship the startup handshake independently. This issue is not a prerequisite for fixing the observed slow-start duplicate batch.
Why this is separate
A pending invocation prevents cron from repeatedly requesting startup while the previous request remains unresolved. It does not prove a previously running worker died when its heartbeat becomes stale.
Example: worker A pauses or loses database connectivity; its heartbeat expires; replacement B starts; A resumes. Preventing A from claiming new work or performing stale protected database writes needs ownership enforcement beyond startup acknowledgement.
No stale-owner side effects were demonstrated in the production incident documented in #696. This issue defines the stronger requested guarantee, not another confirmed incident or a new 0.17.2 regression.
Guarantee and platform limits
-breplicas serving the same queue, remain independent. Do not impose queue-wide exclusion or introduce a general autoscaler.Design work
Prefer the smallest protocol that proves the chosen invariant. Session-level advisory locks alone are unsuitable for transaction-pooled connections and do not fence stale external effects. No arbitrary replica-count API or separate scheduling service is requested.
Acceptance criteria
Code and related work
Findings derive from the published
@pgflow/edge-worker@0.17.2source; inspect current main before implementation.Immediate fix: #696 — pending startup invocation and acknowledgement. This issue: stronger ownership during failure/replacement. Implementation order: #696 first; this issue starts only after #696 ships, and may reuse its invocation state. #694 concerns a separate processing stall; no shared cause is established.