Skip to content

Darling startup: a store at its connection limit is retried, and the compose container reports unhealthy once collection has stopped (#4733) - #4806

Merged
erikdarlingdata merged 4 commits into
devfrom
fix/4733-startup-53300-and-healthcheck
Sep 29, 2026
Merged

erikdarlingdata merged 4 commits into
devfrom
fix/4733-startup-53300-and-healthcheck

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Fixes #4733

Why

Two gaps in how Darling handles a failure at startup.

  1. A store at its connection limit when the service starts was treated as a permanent failure. StartupFailureTriage.IsRetryable is a default-deny allowlist, and all of class 53 was terminal by a documented rule. A bring-your-own or compose store can be briefly full at start: after a restart storm, while the previous instance's connections are still closing, or while another application holds them. The service then stopped collecting until someone restarted it, and the process stayed up, which is the state Darling keeps retrying the store at start instead of never collecting after the retry budget (#4508) #4509 set out to remove. SQLSTATE 53300 (too many connections) clears by itself when other connections close. The other class 53 states (53100 disk full, 53200 out of memory, 53400 configuration limit exceeded) do not.
  2. A stopped collector still looked like a running container. After a terminal startup verdict the process stays up on purpose. Only the store compose service had a healthcheck, so docker compose ps showed the darling container as up although it never collected. Exiting with a non-zero code is not the answer: under restart: unless-stopped a plain config error would restart in a loop.

What changes

Commit 1, 53300 is retried:

  • StartupFailureTriage.cs: PostgresErrorCodes.TooManyConnections joins s_retryableSqlStates, so all three collection-blocking startup steps retry it. Pacing is unchanged: the existing fast budget, then the sustained delay. The sustained arm still logs its critical line once and a warning per attempt after it, so a store that never frees a slot still says why.
  • The class remarks list 53300 as retryable, and the "Terminal, including the ones that are close calls" paragraph now says class 53 is split state by state, the way class 57 already is, and why. 53100, 53200 and 53400 stay terminal.
  • StartupFailureTriageTests.cs: the 53300 row moves from the terminal theory to the retryable one. New decision-level tests drive StartupFailureTriage.NextAction, the decision the three loops make: a 53300 at attempt 1 is RetryFast, and past the fast budget it is RetrySustained on the sustained delay. 53100, 53200 and 53400 are Stop at attempt 1 and after the budget.

Commit 2, the healthcheck:

  • CollectorRuntimeState takes an optional marker path, read once from the DARLING_STOPPED_MARKER environment variable in Program.cs, where the singleton is registered. With no variable set nothing changes: no file is written or removed and nothing is logged. That covers the Windows service and every non-compose deployment.
  • With a path set: a leftover marker is removed when the state object is created (a restarted container keeps its /tmp). The marker is written from PublishStoppedCore, the one path every route to Stopped shares (an exception-path stand-down, a rejected config, and managed mode on a non-Windows machine). It is removed again if the phase ever leaves Stopped through PublishRetrying or PublishCollecting; no caller does that today, so this is a guard. The file holds the UTC time and the snapshot's Detail, the text /api/ping already serves. A fault writing or removing the file is logged once at Warning and never throws.
  • Darling/compose/docker-compose.yml, darling service: DARLING_STOPPED_MARKER: /tmp/darling-stopped in a new environment: block, and a healthcheck ["CMD-SHELL", "test ! -e /tmp/darling-stopped"] with interval: 30s, timeout: 5s, retries: 1, start_period: 30s. It tests a file because the image (mcr.microsoft.com/dotnet/aspnet:10.0, Debian) has /bin/sh and no curl, so it cannot ask /api/ping. A comment above it says what it reports and that restart: unless-stopped does not act on health, so a stopped service stays up and shows unhealthy instead of restarting in a loop.
  • Darling/README.md, compose section: one short paragraph saying the container reports unhealthy after a startup failure that stops collection. Its logs say why, and so does /api/ping when the web dashboard is enabled. It needs the current docker-compose.yml, because the healthcheck and DARLING_STOPPED_MARKER are set there and not in the image.
  • Darling/README.md, the "Cannot reach or migrate the Postgres store" entry: a store at its connection limit (SQLSTATE 53300) joins the transient failures that are retried for two minutes first.
  • Tests: CollectorStoppedMarkerTests (a Stopped verdict writes the file with the time and detail, on all three routes; Retrying and Collecting write nothing; a new state object removes a leftover file; leaving Stopped removes it; no path means no file and an unrelated file is left alone; a write fault and a remove fault are each logged once and never throw; Program.cs reads the variable once and hands it to the state). ComposeDarlingHealthcheckTests reads the real compose file (the fixture CiComposeWorkerSizingTests already uses) and checks the darling service sets DARLING_STOPPED_MARKER and its healthcheck tests the same path, plus an injected-drift copy so a parse that matched nothing cannot pass.

Test plan

  • Part 1 RED on dev's sources: StartupFailureTriageTests 78 total, 3 failed: RetryableSqlStates_AreRetryable(53300) ("53300 (too many connections) must be retryable"), NextAction_TooManyConnectionsAtStart_IsRetryFast (expected RetryFast, actual Stop) and NextAction_TooManyConnectionsPastTheFastBudget_IsRetrySustained (expected RetrySustained, actual Stop). The 53100, 53200 and 53400 Stop cases pass before and after, which is the boundary they pin.
  • Part 1 GREEN: StartupFailureTriageTests, CollectorRuntimeStateTests, DarlingManagedPostgresTests, DarlingStoreUpgradeTests, RepoFileAdoptionTests, RepoFileResolutionEquivalenceTests: 468 total, 0 failed, 21 skipped (live tests that need a store).
  • Part 2 RED on dev's behaviour: CollectorStoppedMarkerTests and ComposeDarlingHealthcheckTests, 12 total, 10 failed. The other two are the negatives that hold on dev (RetryingAndCollecting_WriteNoMarker, WithNoPath_NothingIsWrittenOrRemoved). The RED run used a constructor overload that ignored its arguments so the tests compile; on dev itself the tests do not compile, because the constructor does not exist.
  • Part 2 GREEN, with CollectorRuntimeStateTests, CiComposeWorkerSizingTests and StartupFailureTriageTests: 111 total, 0 failed.
  • Every class that names CollectorRuntimeState or reads the compose file, plus the two above: 304 total, 0 failed, 11 skipped (live tests that need a store).
  • Build of Darling.Tests: 0 warnings, 0 errors. origin/dev had not moved since the branch was cut, so the merge was a no-op.
  • Full Darling.Tests suite, once, without DARLING_TEST_PG: 17376 total, 1 failed, 1140 skipped, 1 not run (the runner does not name it). The failure was RepoFileAdoptionTests.TheLfReadingPins_AreExactlyTheOnesDeclaredHere: my new pin called the LF reader, which that census counts. Fixed by reading Program.cs with the raw reader (the pin's anchors are single-line or cross a break with \s*, so the census needs no change). After the fix I re-ran RepoFileAdoptionTests, RepoFileResolutionEquivalenceTests, CollectorStoppedMarkerTests and ComposeDarlingHealthcheckTests: 17 total, 0 failed.
  • The compose file parses as YAML and the darling service shows the new environment and healthcheck.
  • The full suite was not run a second time after the one-line reader fix; only the classes above were.
  • The live tests (they need DARLING_TEST_PG) were not run.
  • docker compose up with the healthcheck was not run: there is no Docker on this machine. Worth one pass by someone who has it: stop collection with a bad darling.json, and docker compose ps should show darling as unhealthy after about a minute (30s start period, then a 30s interval), and healthy again after fixing the config and restarting.

CHANGELOG

SECTION: Fixed
ENTRY:

…f stopping collection (#4733)

SQLSTATE 53300 (too many connections) is now retryable at the three collection-blocking startup steps. It clears by itself when other connections close, and a bring-your-own or compose store can be briefly full at start after a restart storm or while another application holds its connections. 53100, 53200 and 53400 stay terminal because they need an operator. The retry pacing is unchanged.
…re has stopped collection (#4733)

After a terminal startup verdict the process stays up on purpose, so under compose the darling container showed as running although it never collected. CollectorRuntimeState now takes an optional marker path (DARLING_STOPPED_MARKER, read once in Program.cs) and keeps that file present exactly while the phase is Stopped, holding the published detail and the UTC time. The compose darling service sets the variable and adds a healthcheck that fails while the file exists. Restart: unless-stopped does not act on health, so a stopped service stays up and shows unhealthy instead of restarting in a loop. With no variable set nothing changes. A fault writing or removing the file is logged once at Warning and never throws.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 29, 2026 12:09
@erikdarlingdata
erikdarlingdata merged commit bdd6eab into dev Sep 29, 2026
16 of 17 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/4733-startup-53300-and-healthcheck branch September 29, 2026 13:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant