You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
A Darling server connected for more than 30 days loses its config facts: on-load snapshots age out and nothing re-collects them #3930
This one comes from reading the code. It hasn't been seen in the field.
The on-load config snapshots are collected once per connect, but retention purges them by age. So a Darling service that stays connected to a server for more than 30 days loses that server's only config capture. After that, CONFIG_*, CONFIG_MIN_MAX_MEMORY_NARROW, DB_CONFIG, TRACE_FLAGS and the config audit have nothing to read until the next reconnect.
Why
The schedule.PerformanceMonitor.Collectors/CollectorScheduleDefaults.cs sets server_config, database_config, database_scoped_config and trace_flags to new(0, 30): frequency 0 (on-load only) with 30-day retention. server_properties is new(0, 365).
Collection.DarlingWorker's connect path runs frequency-0 collectors once per connect ("On-load config snapshots (effective FrequencyMinutes 0) run once per connect"). Nothing re-runs them on a healthy, long-lived connection.
Retention.DarlingRetention drops chunks (or DELETEs rows) older than the collector's resolved retention, measured on capture_time. EffectivePurgeRetentionDays has no exemption for a table's newest capture.
On a stable production box, 30 days without a reconnect is ordinary. At day 31 the config facts drop out, and so do the config audit's recommendations. They come back only after the next restart or connection failure.
Lite collects on the same schedule but doesn't purge per collector. Hot rows move to parquet, and RetentionService deletes archive files after ArchiveRetentionMonths (3). So on Lite the gap only opens after three months without a restart.
Not folded into #3980: it needs a ruling on the mechanism, not a change to a read. Re-collecting the on-load collectors on a slow cadence changes the scheduler in all three SKUs and adds a capture a day to database_config and the other snapshot tables, which the change-feed tools then diff. Exempting each server's newest capture from the purge works against drop_chunks, which drops whole chunks. Raising retention only postpones the gap. #3929's cheapest fix waits on this choice.
Ruled (Erik, 2026-09-23): re-capture the on-load config collectors daily, as well as on connect, so a server connected for more than 30 days keeps a snapshot inside retention. Same fix as #3929. Worked in tonight's wave.
Summary
This one comes from reading the code. It hasn't been seen in the field.
The on-load config snapshots are collected once per connect, but retention purges them by age. So a Darling service that stays connected to a server for more than 30 days loses that server's only config capture. After that,
CONFIG_*,CONFIG_MIN_MAX_MEMORY_NARROW,DB_CONFIG,TRACE_FLAGSand the config audit have nothing to read until the next reconnect.Why
PerformanceMonitor.Collectors/CollectorScheduleDefaults.cssetsserver_config,database_config,database_scoped_configandtrace_flagstonew(0, 30): frequency 0 (on-load only) with 30-day retention.server_propertiesisnew(0, 365).DarlingWorker's connect path runs frequency-0 collectors once per connect ("On-load config snapshots (effective FrequencyMinutes 0) run once per connect"). Nothing re-runs them on a healthy, long-lived connection.DarlingRetentiondrops chunks (or DELETEs rows) older than the collector's resolved retention, measured oncapture_time.EffectivePurgeRetentionDayshas no exemption for a table's newest capture.On a stable production box, 30 days without a reconnect is ordinary. At day 31 the config facts drop out, and so do the config audit's recommendations. They come back only after the next restart or connection failure.
Lite collects on the same schedule but doesn't purge per collector. Hot rows move to parquet, and
RetentionServicedeletes archive files afterArchiveRetentionMonths(3). So on Lite the gap only opens after three months without a restart.Options
query_store_health.Related: #3896 (latest-value reads), which is why the on-load reads were left unbounded there.