Repository navigation
Darling keeps retrying a slow managed PostgreSQL start instead of never collecting (#4508) - #4538
Merged
Merged
Conversation
…e retry budget (#4508) After the two-minute retry budget on a retryable store connect/migrate failure runs out, the collection loop now keeps retrying every 60 seconds instead of falling through to the terminal stand-down. A non-retryable failure still stops immediately, exactly as before. The retry decision (fast retry, sustained retry, or stop) is now a pure method, StartupFailureTriage.NextAction, so it can be pinned without a running store.
… and config load (#4508)
…three sites (#4508) The store site's two sustained-retry census pins sliced the worker's source starting from the shared catch filter, which all three sustained arms (store, config, bootstrap) now share verbatim — the first-match slice silently retargeted the config site. Both now anchor on their own site's NextAction(attempt, <site>RetryBudget.Elapsed, ex) call, which is unique per site, and the same pins are added for the config and bootstrap arms. Also adds a pin for CollectorRuntimeState.AttemptStatusText, confirming the fast form renders "attempt N of M" and the sustained form never renders a spent fast-budget cap.
…tained-retry # Conflicts: # Darling/Darling.Tests/CollectorRuntimeStateTests.cs # Darling/Darling.Tests/StartupFailureTriageTests.cs # Darling/PerformanceMonitor.Darling.Service/DarlingWorker.cs
The Retrying arm of DescribePing passed the collector's attempt cap straight into the ping body, so a sustained retry answered attempts: 0 instead of omitting the field \u2014 a cap of zero is as wrong as the earlier attempt 30 of 25. DescribePing now omits attempts once the snapshot reports Sustained. AttemptStatusText was unused by any product code, so it and its test are removed; the doc comment that pointed at it now points at DescribePing.
GET /api/ping now carries sustained and retryEverySeconds while a startup step has spent its fast retry budget and moved to the unbounded 60-second retry; both are omitted otherwise. The published detail text gets the same information appended, so the MCP/Viewer readers of that field and the ping body agree.
erikdarlingdata
marked this pull request as ready for review
September 28, 2026 01:49
erikdarlingdata
added a commit
that referenced
this pull request
Sep 28, 2026
…00Z UTC (#4621) Moves the CHANGELOG entries carried in merged pull-request descriptions into [Unreleased]. The cut is PRs merged after 2026-09-26T17:37:33Z and at or before 2026-09-28T17:40:00Z; the next splice starts after it. - 76 PRs are spliced: Fixed 36, Changed 28 and Added 12, each counted once under its first section. That is 79 bullets: #4481's Fixed entry holds three, and #4548 adds a second group under Added. - 21 PRs have no user-visible entry (None, test-only, CI-only, or deferred to a parent). - The [#4509] link definition, which #4538 also carries, is defined once. - The IMPORTANT upgrade notes from #4489, #4495, #4501, #4506 and #4541 are held for the release cut and are not in this change. - Only CHANGELOG.md changes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #4508
Follows #4509, which gave the store connect/migrate step a sustained retry. This PR gives the managed PostgreSQL start and config load the same retry, and fixes what
/api/pingreports while it runs.Why
The managed PostgreSQL start loop and config load had the same shape #4509 fixed for the store:
A bundled PostgreSQL that takes longer than two minutes to start (a slow disk, a long crash recovery) therefore left the service running for the rest of its life without collecting.
What changes
DarlingWorker.cs) gets the same sustained arm as the store:Retrying, neverStopped;SustainedRetryDelay(60 s), throughStartupFailureTriage.NextAction.CollectorRuntimeState.SnapshotgainsSustained, andPublishRetryingtakes asustainedflag.GET /api/pingdropsattemptsand adds"sustained": trueand"retryEverySeconds"(taken fromSustainedRetryDelay, 60 today). It keepsattempt,step,detailand the 503degradedstatus.SustainedRetryDelayandRetryBudgetby one helper, so/api/ping, MCP and the Viewer show the same text.sustainedorretryEverySecondskey, andattemptsstill shows the cap.Test plan
StartupFailureTriageTests) cover all three sites, each anchored on that site's ownNextAction(attempt, <site budget>.Elapsed, ex)line:Retrying, neverStopped;DescribePing_OmitsAttemptsOnceSustained_ButKeepsItWhileFast(CollectorRuntimeStateTests):degraded,attempt30, noattemptskey,"sustained":true,"retryEverySeconds":60, and the detail phrase;"attempts":25and has neither new key nor the phrase;PublishRetrying_Sustained_AppendsRetryPhraseToDetailpins the detail phrase at the publish site;Assert.Null() Failure … Expected: null / Actual: 0.StartupFailureTriageTests+CollectorRuntimeStateTests+DocCommentHygieneTests:Total: 163, Failed: 0.Lite.Testsisn't touched.CHANGELOG
SECTION: Fixed
ENTRY: - A managed PostgreSQL or a config load that takes more than two minutes to start no longer leaves Darling running without collecting ([#4538]) - The sustained start-up retry from [#4509] now also covers the bundled PostgreSQL's own start and, for failures already treated as transient, loading
darling.json. While that retry runs,/api/pingand the status detail say it is a sustained retry, with the interval, instead of reporting an attempt count past the cap.REF: [#4538]: #4538
REF: [#4509]: #4509