Repository navigation
Darling startup: a store at its connection limit is retried, and the compose container reports unhealthy once collection has stopped (#4733) - #4806
Merged
Conversation
…f stopping collection (#4733) SQLSTATE 53300 (too many connections) is now retryable at the three collection-blocking startup steps. It clears by itself when other connections close, and a bring-your-own or compose store can be briefly full at start after a restart storm or while another application holds its connections. 53100, 53200 and 53400 stay terminal because they need an operator. The retry pacing is unchanged.
…re has stopped collection (#4733) After a terminal startup verdict the process stays up on purpose, so under compose the darling container showed as running although it never collected. CollectorRuntimeState now takes an optional marker path (DARLING_STOPPED_MARKER, read once in Program.cs) and keeps that file present exactly while the phase is Stopped, holding the published detail and the UTC time. The compose darling service sets the variable and adds a healthcheck that fails while the file exists. Restart: unless-stopped does not act on health, so a stopped service stays up and shows unhealthy instead of restarting in a loop. With no variable set nothing changes. A fault writing or removing the file is logged once at Warning and never throws.
erikdarlingdata
marked this pull request as ready for review
September 29, 2026 12:09
…n the web dashboard is on, and a store at its connection limit is retried at start (#4733)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #4733
Why
Two gaps in how Darling handles a failure at startup.
StartupFailureTriage.IsRetryableis a default-deny allowlist, and all of class53was terminal by a documented rule. A bring-your-own or compose store can be briefly full at start: after a restart storm, while the previous instance's connections are still closing, or while another application holds them. The service then stopped collecting until someone restarted it, and the process stayed up, which is the state Darling keeps retrying the store at start instead of never collecting after the retry budget (#4508) #4509 set out to remove. SQLSTATE53300(too many connections) clears by itself when other connections close. The other class53states (53100disk full,53200out of memory,53400configuration limit exceeded) do not.storecompose service had a healthcheck, sodocker compose psshowed thedarlingcontainer as up although it never collected. Exiting with a non-zero code is not the answer: underrestart: unless-stoppeda plain config error would restart in a loop.What changes
Commit 1,
53300is retried:StartupFailureTriage.cs:PostgresErrorCodes.TooManyConnectionsjoinss_retryableSqlStates, so all three collection-blocking startup steps retry it. Pacing is unchanged: the existing fast budget, then the sustained delay. The sustained arm still logs its critical line once and a warning per attempt after it, so a store that never frees a slot still says why.53300as retryable, and the "Terminal, including the ones that are close calls" paragraph now says class53is split state by state, the way class57already is, and why.53100,53200and53400stay terminal.StartupFailureTriageTests.cs: the53300row moves from the terminal theory to the retryable one. New decision-level tests driveStartupFailureTriage.NextAction, the decision the three loops make: a53300at attempt 1 isRetryFast, and past the fast budget it isRetrySustainedon the sustained delay.53100,53200and53400areStopat attempt 1 and after the budget.Commit 2, the healthcheck:
CollectorRuntimeStatetakes an optional marker path, read once from theDARLING_STOPPED_MARKERenvironment variable inProgram.cs, where the singleton is registered. With no variable set nothing changes: no file is written or removed and nothing is logged. That covers the Windows service and every non-compose deployment./tmp). The marker is written fromPublishStoppedCore, the one path every route toStoppedshares (an exception-path stand-down, a rejected config, and managed mode on a non-Windows machine). It is removed again if the phase ever leavesStoppedthroughPublishRetryingorPublishCollecting; no caller does that today, so this is a guard. The file holds the UTC time and the snapshot'sDetail, the text/api/pingalready serves. A fault writing or removing the file is logged once at Warning and never throws.Darling/compose/docker-compose.yml,darlingservice:DARLING_STOPPED_MARKER: /tmp/darling-stoppedin a newenvironment:block, and a healthcheck["CMD-SHELL", "test ! -e /tmp/darling-stopped"]withinterval: 30s,timeout: 5s,retries: 1,start_period: 30s. It tests a file because the image (mcr.microsoft.com/dotnet/aspnet:10.0, Debian) has/bin/shand nocurl, so it cannot ask/api/ping. A comment above it says what it reports and thatrestart: unless-stoppeddoes not act on health, so a stopped service stays up and shows unhealthy instead of restarting in a loop.Darling/README.md, compose section: one short paragraph saying the container reports unhealthy after a startup failure that stops collection. Its logs say why, and so does/api/pingwhen the web dashboard is enabled. It needs the currentdocker-compose.yml, because the healthcheck andDARLING_STOPPED_MARKERare set there and not in the image.Darling/README.md, the "Cannot reach or migrate the Postgres store" entry: a store at its connection limit (SQLSTATE 53300) joins the transient failures that are retried for two minutes first.CollectorStoppedMarkerTests(a Stopped verdict writes the file with the time and detail, on all three routes; Retrying and Collecting write nothing; a new state object removes a leftover file; leaving Stopped removes it; no path means no file and an unrelated file is left alone; a write fault and a remove fault are each logged once and never throw;Program.csreads the variable once and hands it to the state).ComposeDarlingHealthcheckTestsreads the real compose file (the fixtureCiComposeWorkerSizingTestsalready uses) and checks thedarlingservice setsDARLING_STOPPED_MARKERand its healthcheck tests the same path, plus an injected-drift copy so a parse that matched nothing cannot pass.Test plan
StartupFailureTriageTests78 total, 3 failed:RetryableSqlStates_AreRetryable(53300)("53300 (too many connections) must be retryable"),NextAction_TooManyConnectionsAtStart_IsRetryFast(expectedRetryFast, actualStop) andNextAction_TooManyConnectionsPastTheFastBudget_IsRetrySustained(expectedRetrySustained, actualStop). The53100,53200and53400Stopcases pass before and after, which is the boundary they pin.StartupFailureTriageTests,CollectorRuntimeStateTests,DarlingManagedPostgresTests,DarlingStoreUpgradeTests,RepoFileAdoptionTests,RepoFileResolutionEquivalenceTests: 468 total, 0 failed, 21 skipped (live tests that need a store).CollectorStoppedMarkerTestsandComposeDarlingHealthcheckTests, 12 total, 10 failed. The other two are the negatives that hold on dev (RetryingAndCollecting_WriteNoMarker,WithNoPath_NothingIsWrittenOrRemoved). The RED run used a constructor overload that ignored its arguments so the tests compile; on dev itself the tests do not compile, because the constructor does not exist.CollectorRuntimeStateTests,CiComposeWorkerSizingTestsandStartupFailureTriageTests: 111 total, 0 failed.CollectorRuntimeStateor reads the compose file, plus the two above: 304 total, 0 failed, 11 skipped (live tests that need a store).Darling.Tests: 0 warnings, 0 errors.origin/devhad not moved since the branch was cut, so the merge was a no-op.Darling.Testssuite, once, withoutDARLING_TEST_PG: 17376 total, 1 failed, 1140 skipped, 1 not run (the runner does not name it). The failure wasRepoFileAdoptionTests.TheLfReadingPins_AreExactlyTheOnesDeclaredHere: my new pin called the LF reader, which that census counts. Fixed by readingProgram.cswith the raw reader (the pin's anchors are single-line or cross a break with\s*, so the census needs no change). After the fix I re-ranRepoFileAdoptionTests,RepoFileResolutionEquivalenceTests,CollectorStoppedMarkerTestsandComposeDarlingHealthcheckTests: 17 total, 0 failed.darlingservice shows the newenvironmentandhealthcheck.DARLING_TEST_PG) were not run.docker compose upwith the healthcheck was not run: there is no Docker on this machine. Worth one pass by someone who has it: stop collection with a baddarling.json, anddocker compose psshould showdarlingas unhealthy after about a minute (30s start period, then a 30s interval), and healthy again after fixing the config and restarting.CHANGELOG
SECTION: Fixed
ENTRY:
darlingcontainer reports unhealthy when a startup failure has stopped collection ([Darling startup: a store at its connection limit is retried, and the compose container reports unhealthy once collection has stopped (#4733) #4806]) - It used to show as running although it never collected.REF:
[Darling startup: a store at its connection limit is retried, and the compose container reports unhealthy once collection has stopped (#4733) #4806]: Darling startup: a store at its connection limit is retried, and the compose container reports unhealthy once collection has stopped (#4733) #4806