Skip to content

Darling keeps retrying a slow managed PostgreSQL start instead of never collecting (#4508) - #4538

Merged
erikdarlingdata merged 6 commits into
devfrom
fix/4508-bootstrap-sustained-retry
Sep 28, 2026
Merged

erikdarlingdata merged 6 commits into
devfrom
fix/4508-bootstrap-sustained-retry

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Refs #4508

Follows #4509, which gave the store connect/migrate step a sustained retry. This PR gives the managed PostgreSQL start and config load the same retry, and fixes what /api/ping reports while it runs.

Why

The managed PostgreSQL start loop and config load had the same shape #4509 fixed for the store:

  • two minutes of fast retries on a failure the classifier calls retryable;
  • then a fall-through to the terminal catch that a non-retryable failure takes.

A bundled PostgreSQL that takes longer than two minutes to start (a slow disk, a long crash recovery) therefore left the service running for the rest of its life without collecting.

What changes

  • The managed PostgreSQL start (DarlingWorker.cs) gets the same sustained arm as the store:
    • once the fast budget is spent on a still-retryable failure, it logs CRITICAL once, then a WARNING on each retry;
    • it publishes Retrying, never Stopped;
    • it retries every SustainedRetryDelay (60 s), through StartupFailureTriage.NextAction.
  • Config load gets the identical arm, only for failures it already classified as retryable. No new failure class becomes retryable, and a non-retryable failure still stops.
  • The sustained flag: CollectorRuntimeState.Snapshot gains Sustained, and PublishRetrying takes a sustained flag.
  • The status says it's a sustained retry. Before, it reported an attempt number past the fast budget's cap ("attempt 30 of 25"). Now, once the retry is sustained:
    • GET /api/ping drops attempts and adds "sustained": true and "retryEverySeconds" (taken from SustainedRetryDelay, 60 today). It keeps attempt, step, detail and the 503 degraded status.
    • The step detail ends with "retrying every 60s, attempt N (past the 120s fast budget)", built from SustainedRetryDelay and RetryBudget by one helper, so /api/ping, MCP and the Viewer show the same text.
    • A fast retry is unchanged: no sustained or retryEverySeconds key, and attempts still shows the cap.

Test plan

  • The code-reading checks (StartupFailureTriageTests) cover all three sites, each anchored on that site's own NextAction(attempt, <site budget>.Elapsed, ex) line:
    • CRITICAL once, then a WARNING per retry;
    • publishes Retrying, never Stopped;
    • the store site's two existing checks now use the same per-site anchor, because all three sites share one catch filter.
  • DescribePing_OmitsAttemptsOnceSustained_ButKeepsItWhileFast (CollectorRuntimeStateTests):
    • a sustained snapshot gives 503, degraded, attempt 30, no attempts key, "sustained":true, "retryEverySeconds":60, and the detail phrase;
    • a fast snapshot keeps "attempts":25 and has neither new key nor the phrase;
    • PublishRetrying_Sustained_AppendsRetryPhraseToDetail pins the detail phrase at the publish site;
    • RED at runtime before the ping change: Assert.Null() Failure … Expected: null / Actual: 0.
  • Mutations, each reverted:
    • removing the managed PostgreSQL start's sustained arm fails its two new checks;
    • reverting the ping change fails the ping test.
  • The run: StartupFailureTriageTests + CollectorRuntimeStateTests + DocCommentHygieneTests: Total: 163, Failed: 0.
  • Windows-only: Lite.Tests isn't touched.

CHANGELOG

SECTION: Fixed
ENTRY: - A managed PostgreSQL or a config load that takes more than two minutes to start no longer leaves Darling running without collecting ([#4538]) - The sustained start-up retry from [#4509] now also covers the bundled PostgreSQL's own start and, for failures already treated as transient, loading darling.json. While that retry runs, /api/ping and the status detail say it is a sustained retry, with the interval, instead of reporting an attempt count past the cap.
REF: [#4538]: #4538
REF: [#4509]: #4509

…e retry budget (#4508)

After the two-minute retry budget on a retryable store connect/migrate
failure runs out, the collection loop now keeps retrying every 60
seconds instead of falling through to the terminal stand-down. A
non-retryable failure still stops immediately, exactly as before.

The retry decision (fast retry, sustained retry, or stop) is now a
pure method, StartupFailureTriage.NextAction, so it can be pinned
without a running store.
…three sites (#4508)

The store site's two sustained-retry census pins sliced the worker's source starting from the shared catch filter, which all three sustained arms (store, config, bootstrap) now share verbatim — the first-match slice silently retargeted the config site. Both now anchor on their own site's NextAction(attempt, <site>RetryBudget.Elapsed, ex) call, which is unique per site, and the same pins are added for the config and bootstrap arms. Also adds a pin for CollectorRuntimeState.AttemptStatusText, confirming the fast form renders "attempt N of M" and the sustained form never renders a spent fast-budget cap.
…tained-retry

# Conflicts:
#	Darling/Darling.Tests/CollectorRuntimeStateTests.cs
#	Darling/Darling.Tests/StartupFailureTriageTests.cs
#	Darling/PerformanceMonitor.Darling.Service/DarlingWorker.cs
The Retrying arm of DescribePing passed the collector's attempt cap straight
into the ping body, so a sustained retry answered attempts: 0 instead of
omitting the field \u2014 a cap of zero is as wrong as the earlier attempt 30 of
25. DescribePing now omits attempts once the snapshot reports Sustained.

AttemptStatusText was unused by any product code, so it and its test are
removed; the doc comment that pointed at it now points at DescribePing.
GET /api/ping now carries sustained and retryEverySeconds while a
startup step has spent its fast retry budget and moved to the
unbounded 60-second retry; both are omitted otherwise. The published
detail text gets the same information appended, so the MCP/Viewer
readers of that field and the ping body agree.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 28, 2026 01:49
@erikdarlingdata
erikdarlingdata merged commit 0524634 into dev Sep 28, 2026
17 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/4508-bootstrap-sustained-retry branch September 28, 2026 01:49
erikdarlingdata added a commit that referenced this pull request Sep 28, 2026
…00Z UTC (#4621)

Moves the CHANGELOG entries carried in merged pull-request descriptions into [Unreleased]. The cut is PRs merged after 2026-09-26T17:37:33Z and at or before 2026-09-28T17:40:00Z; the next splice starts after it.

- 76 PRs are spliced: Fixed 36, Changed 28 and Added 12, each counted once under its first section. That is 79 bullets: #4481's Fixed entry holds three, and #4548 adds a second group under Added.
- 21 PRs have no user-visible entry (None, test-only, CI-only, or deferred to a parent).
- The [#4509] link definition, which #4538 also carries, is defined once.
- The IMPORTANT upgrade notes from #4489, #4495, #4501, #4506 and #4541 are held for the release cut and are not in this change.
- Only CHANGELOG.md changes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant