Skip to content

Latest-value lookbacks follow the collector's cadence: a day, or twice its interval if longer (#3896) - #3980

Merged
erikdarlingdata merged 2 commits into
devfrom
feature/3896-cadence-lookback
Sep 23, 2026
Merged

erikdarlingdata merged 2 commits into
devfrom
feature/3896-cadence-lookback

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Closes #3896.

Follow-up to #3931, which merged the lookback as a flat day. This PR adds the cadence-aware lookback requested in #3931's review. It also carries #3928 as a second commit, under the rule that small in-lane follow-ups go in the PR: the PostgreSQL target's config reads get the same planning fix without changing a single answer.

Why

#3931 bounds each analysis latest-value read to one day before the window's end. It covers the database-size, percent-autogrowth, disk-space, memory-clerk, plan-cache and memory_stats facts, and the autogrowth drill-down. But operators can edit collector cadences: Darling's config_collector_schedules.frequency_minutes has no upper bound, and Lite and the Dashboard keep editable schedules too. A steady-state collection advances by exactly its interval (#1553). So a database_size_stats, memory_clerks or plan_cache_stats collector scheduled daily or slower loses its facts on every pass between 24 h after its last sample and its next run. That's a one-pass flap, and it resolves and then re-fires the findings and notifications built on those facts.

#3928: the PostgreSQL target's five reads of the newest pg_server_config snapshot had no lower bound at all. The table keeps a year of hourly snapshots in one-day chunks, so a year in, TimescaleDB planned every chunk on every read: 4.0-8.3 s of planning per read at 366 chunks on the rig. Four of the reads run on every analysis pass, and the clock read runs once for every baseline the pass computes.

What changes

  • The lookback is a day, or twice the collector's interval if that is longer (AnalysisContext.LatestValueLookbackFor). Twice the interval keeps the newest sample in range between runs, with a whole interval to spare for a late or failed one. Every shipped cadence is under twelve hours, so default behavior is unchanged: the flat day, as Analysis latest-value reads look back a day instead of scanning a server's whole history, and stop counting dropped databases (#3896) #3931 shipped it.
  • The cadence is the one the collector actually runs at. It's resolved once per pass and stamped on AnalysisContext.LatestValueStarts, and every read binds LatestValueStartFor(<its collector>). The drill-down reads the same stamp, so a list and the count it details stay bounded alike.
    • Darling reads config_collector_schedules (per-server over fleet-wide) through the rule the scheduler runs by. That rule moved from StoreConfigProvider to CollectorScheduleDefaults.ResolveFrequencyMinutes, and ResolveSchedule delegates to it, so the scheduler and the analysis can't resolve two different cadences. If the overrides can't be read, the pass falls back to the defaults and logs a warning.
    • Lite resolves its ScheduleManager through the storage-hash to connection-GUID match the alert adapter already used. That match is now one method, ScheduleManager.GetFrequencyForStorageServer. AnalysisService takes the resolver, and the scheduled pass, the MCP host and the Recommendations tab all pass it. With no resolver, every collector takes its shipped default.
    • Dashboard reads config.collection_schedule. Its frequency 0 means "every tick", so it takes the day.
  • On-load collectors. A collector moved to on-load only (frequency 0) anchors on its newest capture at or before the window's end, however old, so it doesn't lose the fact a day after every connect.
  • The bound stays a plain timestamp parameter, computed in C#, so TimescaleDB still excludes chunks at plan time at any lookback.
  • No SQL text changes for these reads: they keep Analysis latest-value reads look back a day instead of scanning a server's whole history, and stop counting dropped databases (#3896) #3931's collection_time >= $2 AND collection_time <= $3, and the Lite twins stay byte-identical.

#3928: the PostgreSQL target's config reads (Darling only; Lite and the Dashboard don't monitor PostgreSQL)

  • The config, posture, blocking-settings and memory families' reads and the baseline clock's (PgTargetBaselineProvider.PgTargetClockSql) bound both the snapshot's MAX(collection_time) and the row scan from below. Bounding only the MAX isn't enough, because the outer collection_time = (…) can exclude chunks only at run time.
  • Every read runs the day first, and unbounded only when that found nothing (PgTargetFactCollector.ConfigSnapshotLowerBounds: the day, then -infinity). The newest snapshot in the day is the newest snapshot the unbounded anchor finds, so the two runs can't disagree:
    • a target whose config collector has been dark for days still states its newest snapshot, with its age on the fact;
    • a historical window still reads the snapshot at or before its own end;
    • posture still reads the newest snapshot the target has.
  • That removes the ruling PostgreSQL-target config reads plan every retained pg_server_config chunk: seconds per read once a store holds a year of snapshots #3928 asked for. Posture's "the snapshot, not the window" rule and the clock's UTC fallback behave exactly as before.
  • These reads take a flat day plus the fallback, not their collector's cadence. They anchor on one snapshot per server, so there's no stale series for a bound to drop, and a cadence slower than a day costs the fallback run, never a fact.

Test plan

Left open

🤖 Generated with Claude Code

https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif

erikdarlingdata and others added 2 commits September 22, 2026 21:24
…e its interval if that is longer (#3896)

Collector frequencies are operator-editable and a steady-state collection
advances by exactly the interval, so a flat 24-hour lookback dropped a daily
collector's facts on every pass between the day and its next run, resolving
and re-firing the findings built on them.

Each latest-value read now binds its collector's lower bound, stamped once
per pass on AnalysisContext.LatestValueStarts from the EFFECTIVE schedule:
Darling reads config_collector_schedules through the rule the scheduler runs
by (now shared as CollectorScheduleDefaults.ResolveFrequencyMinutes), Lite
resolves ScheduleManager through the storage-hash match the alert adapter
already used, and the Dashboard reads config.collection_schedule. An on-load
collector (frequency 0) anchors on its newest capture. The bound stays a
plain parameter, so TimescaleDB still excludes chunks at plan time: a daily
collector's read plans 3 chunks where the unbounded read planned 11. Every
shipped cadence keeps the flat day.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
…hunks, not a year, and still state the newest snapshot when the collector is dark (#3928)

The five reads of the newest pg_server_config snapshot (the config, posture,
blocking-settings and memory families, and the baseline clock) now bound both
the snapshot's MAX(collection_time) and the row scan from below by the day
before the window's end. Each runs unbounded only when the day found nothing
(PgTargetFactCollector.ConfigSnapshotLowerBounds: the day, then -infinity).

The newest snapshot in the day is the newest snapshot the unbounded anchor
finds, so no answer changes. A target whose config collector is dark still
states its newest snapshot, a historical window still reads its own, and
posture still reads the newest the target has.

At 366 one-day chunks on the rig, planning per read drops from 4.0-8.3 s to
0.6-1.1 ms. On DARLING01 (17 chunks, TimescaleDB 2.30.1) it drops from 7.1-7.6 ms
to 0.8 ms, with the same rows.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
@erikdarlingdata
erikdarlingdata merged commit badacaf into dev Sep 23, 2026
15 of 16 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/3896-cadence-lookback branch September 23, 2026 02:59
erikdarlingdata added a commit that referenced this pull request Sep 23, 2026
#3980 merged to dev with two more ChunkScans callers in
LatestValueLookbackTests and a fourth copy of the double-counting pattern
in PgTargetConfigSnapshotBoundTests. Both now use the shared helper.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv
erikdarlingdata added a commit that referenced this pull request Sep 23, 2026
…ection-health policy run waits out a launch that beat the park (#3982)

* Chunk-exclusion tests count a bitmap plan's chunks once, and the collection-health policy run waits out a launch that beat the park

Two CI flakes that blocked #3968 and #3972 on code neither PR touched.

LatestValueLookbackLivePostgresTests reported "planned 4 chunks" for a
24-hour read whose plan scanned two. TimescaleDB names a chunk's indexes
after the chunk, so under a bitmap plan the "Bitmap Index Scan on
_hyper_12_11_chunk_..._idx" child matched the chunk pattern too, and the
count doubled whenever the planner picked bitmap scans. The same pattern
lived in FleetReadsAreBoundedByTheFleetTests and CaptureDownChunkOrderTests.
One shared PlanChunkScans helper now ends the match on the chunk name, and
PlanChunkScansTests pins it against the plan CI printed.

CollectionHealthAggregateTests' run_job failed 55P03 "concurrent refresh":
collection_health_hourly's policy has no initial_start, so TimescaleDB
launches its first run the moment the product commits it, before the
fixture can park it. RunPolicyAsync now retries 55P03 for up to 30 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

* The chunk counters #3980 added go through PlanChunkScans too

#3980 merged to dev with two more ChunkScans callers in
LatestValueLookbackTests and a fourth copy of the double-counting pattern
in PgTargetConfigSnapshotBoundTests. Both now use the shared helper.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
erikdarlingdata added a commit that referenced this pull request Sep 23, 2026
…3920, #3927, #3931, #3932, #3940, #3942, #3946, #3947, #3950, #3952, #3955, #3956, #3957, #3964, #3965, #3966, #3968, #3972, #3975, #3979, #3980, #3981, #3983, #3984, #3985) (#3989)

The wave's fix PRs deliberately carried no CHANGELOG edits (parallel-agent hot-spot protocol); each agent reported its entry and this commit lands them together, byte-verified against origin/dev.


Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant