Skip to content

A Darling server connected for more than 30 days loses its config facts: on-load snapshots age out and nothing re-collects them #3930

Description

@erikdarlingdata

Summary

This one comes from reading the code. It hasn't been seen in the field.

The on-load config snapshots are collected once per connect, but retention purges them by age. So a Darling service that stays connected to a server for more than 30 days loses that server's only config capture. After that, CONFIG_*, CONFIG_MIN_MAX_MEMORY_NARROW, DB_CONFIG, TRACE_FLAGS and the config audit have nothing to read until the next reconnect.

Why

  • The schedule. PerformanceMonitor.Collectors/CollectorScheduleDefaults.cs sets server_config, database_config, database_scoped_config and trace_flags to new(0, 30): frequency 0 (on-load only) with 30-day retention. server_properties is new(0, 365).
  • Collection. DarlingWorker's connect path runs frequency-0 collectors once per connect ("On-load config snapshots (effective FrequencyMinutes 0) run once per connect"). Nothing re-runs them on a healthy, long-lived connection.
  • Retention. DarlingRetention drops chunks (or DELETEs rows) older than the collector's resolved retention, measured on capture_time. EffectivePurgeRetentionDays has no exemption for a table's newest capture.

On a stable production box, 30 days without a reconnect is ordinary. At day 31 the config facts drop out, and so do the config audit's recommendations. They come back only after the next restart or connection failure.

Lite collects on the same schedule but doesn't purge per collector. Hot rows move to parquet, and RetentionService deletes archive files after ArchiveRetentionMonths (3). So on Lite the gap only opens after three months without a restart.

Options

  1. Re-run the on-load collectors on a slow cadence (daily, say) as well as on connect. Cheapest, and it matches the database_config reports query_store as a boolean — collect its health (actual_state, size vs cap, read-only reason) so a thrashing store is visible #2319 reasoning for query_store_health.
  2. Exempt each server's newest capture from the purge for the frequency-0 tables.
  3. Raise these tables' retention. That only postpones the gap.

Related: #3896 (latest-value reads), which is why the on-load reads were left unbounded there.

Activity

  1. erikdarlingdata commented on Sep 23, 2026

    @erikdarlingdata
    OwnerAuthor

    Not folded into #3980: it needs a ruling on the mechanism, not a change to a read. Re-collecting the on-load collectors on a slow cadence changes the scheduler in all three SKUs and adds a capture a day to database_config and the other snapshot tables, which the change-feed tools then diff. Exempting each server's newest capture from the purge works against drop_chunks, which drops whole chunks. Raising retention only postpones the gap. #3929's cheapest fix waits on this choice.

  2. erikdarlingdata commented on Sep 23, 2026

    @erikdarlingdata
    OwnerAuthor

    Ruled (Erik, 2026-09-23): re-capture the on-load config collectors daily, as well as on connect, so a server connected for more than 30 days keeps a snapshot inside retention. Same fix as #3929. Worked in tonight's wave.

  3. added
    in-progressActively being worked by a local session or its agents (PR open or in flight)
    on Sep 23, 2026
  4. erikdarlingdata commented on Sep 23, 2026

    @erikdarlingdata
    OwnerAuthor

    Closed by the watcher: delivered in PR #4001, merged to dev.

  5. removed
    in-progressActively being worked by a local session or its agents (PR open or in flight)
    on Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions