Repository navigation
Latest-value lookbacks follow the collector's cadence: a day, or twice its interval if longer (#3896) - #3980
Merged
Merged
Conversation
…e its interval if that is longer (#3896) Collector frequencies are operator-editable and a steady-state collection advances by exactly the interval, so a flat 24-hour lookback dropped a daily collector's facts on every pass between the day and its next run, resolving and re-firing the findings built on them. Each latest-value read now binds its collector's lower bound, stamped once per pass on AnalysisContext.LatestValueStarts from the EFFECTIVE schedule: Darling reads config_collector_schedules through the rule the scheduler runs by (now shared as CollectorScheduleDefaults.ResolveFrequencyMinutes), Lite resolves ScheduleManager through the storage-hash match the alert adapter already used, and the Dashboard reads config.collection_schedule. An on-load collector (frequency 0) anchors on its newest capture. The bound stays a plain parameter, so TimescaleDB still excludes chunks at plan time: a daily collector's read plans 3 chunks where the unbounded read planned 11. Every shipped cadence keeps the flat day. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
…hunks, not a year, and still state the newest snapshot when the collector is dark (#3928) The five reads of the newest pg_server_config snapshot (the config, posture, blocking-settings and memory families, and the baseline clock) now bound both the snapshot's MAX(collection_time) and the row scan from below by the day before the window's end. Each runs unbounded only when the day found nothing (PgTargetFactCollector.ConfigSnapshotLowerBounds: the day, then -infinity). The newest snapshot in the day is the newest snapshot the unbounded anchor finds, so no answer changes. A target whose config collector is dark still states its newest snapshot, a historical window still reads its own, and posture still reads the newest the target has. At 366 one-day chunks on the rig, planning per read drops from 4.0-8.3 s to 0.6-1.1 ms. On DARLING01 (17 chunks, TimescaleDB 2.30.1) it drops from 7.1-7.6 ms to 0.8 ms, with the same rows. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
This was referenced Sep 23, 2026
erikdarlingdata
added a commit
that referenced
this pull request
Sep 23, 2026
#3980 merged to dev with two more ChunkScans callers in LatestValueLookbackTests and a fourth copy of the double-counting pattern in PgTargetConfigSnapshotBoundTests. Both now use the shared helper. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv
erikdarlingdata
added a commit
that referenced
this pull request
Sep 23, 2026
…ection-health policy run waits out a launch that beat the park (#3982) * Chunk-exclusion tests count a bitmap plan's chunks once, and the collection-health policy run waits out a launch that beat the park Two CI flakes that blocked #3968 and #3972 on code neither PR touched. LatestValueLookbackLivePostgresTests reported "planned 4 chunks" for a 24-hour read whose plan scanned two. TimescaleDB names a chunk's indexes after the chunk, so under a bitmap plan the "Bitmap Index Scan on _hyper_12_11_chunk_..._idx" child matched the chunk pattern too, and the count doubled whenever the planner picked bitmap scans. The same pattern lived in FleetReadsAreBoundedByTheFleetTests and CaptureDownChunkOrderTests. One shared PlanChunkScans helper now ends the match on the chunk name, and PlanChunkScansTests pins it against the plan CI printed. CollectionHealthAggregateTests' run_job failed 55P03 "concurrent refresh": collection_health_hourly's policy has no initial_start, so TimescaleDB launches its first run the moment the product commits it, before the fixture can park it. RunPolicyAsync now retries 55P03 for up to 30 s. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv * The chunk counters #3980 added go through PlanChunkScans too #3980 merged to dev with two more ChunkScans callers in LatestValueLookbackTests and a fourth copy of the double-counting pattern in PgTargetConfigSnapshotBoundTests. Both now use the shared helper. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv --------- Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
erikdarlingdata
added a commit
that referenced
this pull request
Sep 23, 2026
…3920, #3927, #3931, #3932, #3940, #3942, #3946, #3947, #3950, #3952, #3955, #3956, #3957, #3964, #3965, #3966, #3968, #3972, #3975, #3979, #3980, #3981, #3983, #3984, #3985) (#3989) The wave's fix PRs deliberately carried no CHANGELOG edits (parallel-agent hot-spot protocol); each agent reported its entry and this commit lands them together, byte-verified against origin/dev. Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
5 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3896.
Follow-up to #3931, which merged the lookback as a flat day. This PR adds the cadence-aware lookback requested in #3931's review. It also carries #3928 as a second commit, under the rule that small in-lane follow-ups go in the PR: the PostgreSQL target's config reads get the same planning fix without changing a single answer.
Why
#3931 bounds each analysis latest-value read to one day before the window's end. It covers the database-size, percent-autogrowth, disk-space, memory-clerk, plan-cache and memory_stats facts, and the autogrowth drill-down. But operators can edit collector cadences: Darling's
config_collector_schedules.frequency_minuteshas no upper bound, and Lite and the Dashboard keep editable schedules too. A steady-state collection advances by exactly its interval (#1553). So a database_size_stats, memory_clerks or plan_cache_stats collector scheduled daily or slower loses its facts on every pass between 24 h after its last sample and its next run. That's a one-pass flap, and it resolves and then re-fires the findings and notifications built on those facts.#3928: the PostgreSQL target's five reads of the newest
pg_server_configsnapshot had no lower bound at all. The table keeps a year of hourly snapshots in one-day chunks, so a year in, TimescaleDB planned every chunk on every read: 4.0-8.3 s of planning per read at 366 chunks on the rig. Four of the reads run on every analysis pass, and the clock read runs once for every baseline the pass computes.What changes
AnalysisContext.LatestValueLookbackFor). Twice the interval keeps the newest sample in range between runs, with a whole interval to spare for a late or failed one. Every shipped cadence is under twelve hours, so default behavior is unchanged: the flat day, as Analysis latest-value reads look back a day instead of scanning a server's whole history, and stop counting dropped databases (#3896) #3931 shipped it.AnalysisContext.LatestValueStarts, and every read bindsLatestValueStartFor(<its collector>). The drill-down reads the same stamp, so a list and the count it details stay bounded alike.config_collector_schedules(per-server over fleet-wide) through the rule the scheduler runs by. That rule moved fromStoreConfigProvidertoCollectorScheduleDefaults.ResolveFrequencyMinutes, andResolveScheduledelegates to it, so the scheduler and the analysis can't resolve two different cadences. If the overrides can't be read, the pass falls back to the defaults and logs a warning.ScheduleManagerthrough the storage-hash to connection-GUID match the alert adapter already used. That match is now one method,ScheduleManager.GetFrequencyForStorageServer.AnalysisServicetakes the resolver, and the scheduled pass, the MCP host and the Recommendations tab all pass it. With no resolver, every collector takes its shipped default.config.collection_schedule. Its frequency 0 means "every tick", so it takes the day.collection_time >= $2 AND collection_time <= $3, and the Lite twins stay byte-identical.#3928: the PostgreSQL target's config reads (Darling only; Lite and the Dashboard don't monitor PostgreSQL)
PgTargetBaselineProvider.PgTargetClockSql) bound both the snapshot'sMAX(collection_time)and the row scan from below. Bounding only theMAXisn't enough, because the outercollection_time = (…)can exclude chunks only at run time.PgTargetFactCollector.ConfigSnapshotLowerBounds: the day, then-infinity). The newest snapshot in the day is the newest snapshot the unbounded anchor finds, so the two runs can't disagree:Test plan
LatestValueLookbackTests(53, up from 27):ResolveFrequencyMinutesandStoreConfigProvider.ResolveSchedule, for per-server, fleet, negative and delta-capped overrides;FILE_AUTOGROWTH_PERCENT,DISK_SPACEand the drill-down file, and a 49 h-old one does not, while a server left at the hourly default loses the same 25 h sample exactly as before;LatestValueLookbackTests(6): the daily override (25 h kept, 49 h dropped, default dropped) and the on-load anchor on DuckDB, through the production resolver over a realServerManagerandScheduleManager.LatestValueLookbackSqlTests(9): each read bound by its own collector's stamp, the stamp taken before the first read and in the drill-down, and the arithmetic offconfig.collection_schedule(1,440 → 48 h; 1, 0 and no row → 24 h).config_collector_schedulesrow enableslong_query_completions). Its reads therefore keep Analysis latest-value reads look back a day instead of scanning a server's whole history, and stop counting dropped databases (#3896) #3931's measured bounds and numbers:DatabaseSizeSql1,167 → 56 ms, and 1.31 s → 66 ms per fact pass.PgTargetConfigSnapshotBoundTests(14 new: 12 ungated, 2 live):snapshot_age_s259,200) and its clock zone; a window that ended five days ago reads its own snapshot for config and the clock, and posture reads today's.PgTargetClockTests(the shared anchor text, token for token),PgTargetKnobsTests,PgTargetMemoryTests,PgTargetPostureTests,PgTargetBlockingTests, and the V138 scope census inPgServerConfigScopeRungTests(whose anchor-subquery pattern now allows the lower bound; still 12 reads).PgTargetClockLiveTestsalready crossed the fallback: its New York snapshot is 31 days older than the window's end.ConfigSnapshotLowerBounds, 4 tests fail, including the pre-existing clock e2e.Darling.TestswithDARLING_TEST_PG(rig: PostgreSQL 18.4 + TimescaleDB 2.28.1): 12,732, 0 failed, 24 skipped (the runtime and upgrade tests gated on variables this rig leaves unset). An earlier full run, before the census fix, failed 3: the V138 scope census, fixed in the PostgreSQL-target config reads plan every retained pg_server_config chunk: seconds per read once a store holds a year of snapshots #3928 commit; and two load flakes that pass alone,PgWaitSamplerLiveTests(PgWaitSamplerLiveTests can count another test's lock wait: 31 samples from a 30-snapshot cycle under full-suite load #3939) andStoreToastAndCheckpointerLivePostgresTests(filed as Test flake: StoreToastAndCheckpointerLivePostgresTests reads 96 % live after deleting half the rows under full-suite load #3977).Lite.Tests: 5,134, 0 failed (on the cadence commit; the PostgreSQL-target config reads plan every retained pg_server_config chunk: seconds per read once a store holds a year of snapshots #3928 commit touches no Lite or shared code).Dashboard.Tests: 818, 0 failed (same).Left open
TRACE_FLAGSkeeps a flag after it's turned off) needs a design decision, and the cheapest fix depends on A Darling server connected for more than 30 days loses its config facts: on-load snapshots age out and nothing re-collects them #3930's. If on-load collectors re-run on a slow cadence, the trace-flags read can take this PR's cadence-aware lookback, with no sentinel row and no migration. If they don't, it needs a sentinel row per capture (a change to what everytrace_flagsreader sees) or acollection_loganchor (a new index, so a store migration).get_pg_server_config,get_pg_logging_auditand the viewer's PostgreSQL config panel have the same unbounded anchor. The fallback carries over to two of those statements.OverrideSqlneeds a shape decision first, because it's usually empty, so its bounded run can't tell "no snapshot in the day" from "no overrides".🤖 Generated with Claude Code
https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif