Skip to content

Terminalize tasks after the stalled-task requeue limit #689

Description

@cpursley

Problem

The task can be requeued three times. A worker can then claim it for a fourth attempt and disappear again. On the next recovery sweep, requeue_stalled_tasks() archives its PGMQ message and sets step_tasks.permanently_stalled_at, but leaves step_tasks.status = 'started'. In a one-step flow, step_states.status and runs.status also remain started, with no message left to drive another transition. In a larger flow, dependent steps can remain blocked.

This is the current behavior in the recovery function and its max-requeue test. It means whenExhausted: 'fail', 'skip', or 'skip-cascade' is never applied to this failure mode, even though fail_task() already owns those transitions.

I see that the troubleshooting guide deliberately leaves the status started for investigation. The timestamp and requeue count would still identify the incident if the task became terminal; keeping it active also leaves the flow unable to finish or run a dependent cleanup step.

Suggested direction

Keep permanently_stalled_at as the diagnostic marker, but hand stamped tasks to fail_task() after recovery commits. For each candidate, lock and recheck the run, step, and task; raise attempts_count to at least the step's effective max_attempts; then call fail_task(run_id, step_slug, task_index, 'permanently stalled'). That lets the existing code decide whether to fail the run, skip the step, cascade a skip, start dependents, and cancel or archive siblings. The pass should also pick up already-stamped rows from earlier deployments and be safe to repeat.

I would keep this as a separate database call or scheduled job, rather than calling fail_task() inside requeue_stalled_tasks(). Recovery locks step_tasks first, while fail_task() locks the run and step before updating the task. Holding the recovery locks while entering fail_task() creates a lock-order inversion with live workers. The recovery call should commit before escalation begins.

We implemented and tested this approach in the Elixir port's 0.5.0 commit. The relevant pieces are the V06 SQL function, separate recovery calls, and tests for fail, skip, skip-cascade, map siblings, late completion, and previously stamped rows.

Would you be open to making a permanently stalled task terminal through the existing exhaustion policy while retaining the timestamp for investigation?

Activity

  1. jumski commented on Sep 24, 2026

    @jumski
    Contributor

    Thans for the report @cpursley ! will handle it soon

  2. jumski commented on Oct 1, 2026

    @jumski
    Contributor

    hey @cpursley! thanks for the detailed report and the Elixir example. i dont have an ETA yet, but ill handle this.

    i checked the current SQL and the original recovery change - you are right. We kept tasks started for investigation, but permanently_stalled_at already identifies them independently of status. Keeping the run active is not necessary for that.

    Your proposed direction makes sense. Leaving few implementation notes here so we keep the scope in one place:

    • Keep permanently_stalled_at and requeued_count for diagnosis. Once recovery gives up, apply the existing whenExhausted policy through fail_task() rather than duplicate its task/step/run transitions.
    • Run escalation in a separate transaction after recovery commits. Two function calls inside one transaction are not enough: recovery takes task locks first, while the failure path takes parent locks first. The cron setup needs to preserve that boundary.
    • Lock and recheck the run, step, and task before escalation. Pick up already-stamped rows from older deployments, make repeated sweeps safe, and leave work that already reached a terminal state unchanged.
    • Attempt accounting still needs a decision. Raising attempts_count to the effective max_attempts forces the right branch, but can make the count larger than the actual number of claims. i want to check that trade-off against an explicit forced-exhaustion path before choosing; this part is not settled yet.
    • This is related to the private step queues epic Epic: staged private per-step queues and queue identity #653, but remains a separate lifecycle fix. Recovery and fail_task() already use the task's stored queue identity, so we should reuse that instead of adding routing logic. Separate queues alone cannot unblock a dependent step when its parent stays stalled.

    Checks should cover fail, skip, and skip-cascade; old stamped rows and repeated sweeps; completion before escalation and late callbacks after it; and concurrent recovery/escalation/callbacks. Include map siblings and tasks across private step queues, with correct message archival, counters, and events. Also cover an effective max_attempts higher than the actual claim count so escalation cannot accidentally retry an already-archived message.

    The fix also needs the cron/migration wiring and updated stalled-task docs. No new issue needed - we can track the implementation here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions