Repository navigation
Analysis latest-value reads look back a day instead of scanning a server's whole history, and stop counting dropped databases (#3896) - #3931
Merged
Conversation
…d stop counting dropped databases (#3896) The database-size, percent-autogrowth, disk-space, memory-clerk, plan-cache and memory_stats facts, plus the autogrowth drill-down, took each series' newest row with only an upper time bound (or none). The window function numbered every retained row, and "latest ever" kept dropped databases' files in DATABASE_TOTAL_SIZE_MB, and their ALTER DATABASE in the drill-down, until retention aged them out. Each read now binds AnalysisContext.LatestValueStart (window end minus 24 h) beside the window end, in Darling, Lite and the frozen Dashboard. DB_CONFIG had the same ghost databases; being an on-load snapshot it anchors on the newest capture instead, the rule its own drill-down already followed. DARLING01 (SQL2022, 17 days): DatabaseSizeSql 1,167 ms -> 56 ms (571,576 -> 33,750 rows), 239,068 MB (37 files) -> 182,496 MB (27). MemoryClerkSql 129 ms -> 7 ms. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
erikdarlingdata
deleted the
feature/3896-analysis-latest-value-reads
branch
September 23, 2026 01:23
This was referenced Sep 23, 2026
erikdarlingdata
added a commit
that referenced
this pull request
Sep 23, 2026
…3920, #3927, #3931, #3932, #3940, #3942, #3946, #3947, #3950, #3952, #3955, #3956, #3957, #3964, #3965, #3966, #3968, #3972, #3975, #3979, #3980, #3981, #3983, #3984, #3985) (#3989) The wave's fix PRs deliberately carried no CHANGELOG edits (parallel-agent hot-spot protocol); each agent reported its entry and this commit lands them together, byte-verified against origin/dev. Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3896.
Why
The analysis "latest value" reads took each series' newest row with an upper time bound only, or no bound at all. They are the database-size, percent-autogrowth, disk-space, memory-clerk, plan-cache and memory_stats facts, plus the autogrowth drill-down. The window function numbered every row the server had retained to keep a few dozen, and TimescaleDB could exclude no chunk. On Lite, where the
v_views union the parquet archive, nothing could prune a row group.The answer was also wrong. "Latest ever" kept a dropped database's files in
DATABASE_TOTAL_SIZE_MBuntil retention aged them out. The same files, each with anALTER DATABASEfor a database that no longer exists, stayed in the autogrowth drill-down. Erik's production note stands: on the largest storeDatabaseSizeSqlis about a tenth ofanalyze_server, andPlanRegressionSql(#3902) dominates. This fix is the correctness half plus that tenth.The diagnosis is confirmed and extended by one read.
DB_CONFIGcounted the same five dropped databases on DARLING01: 12 databases per-database-latest against 7 in the newest capture.What changes
AnalysisContext.LatestValueLookbackis 24 h, andLatestValueStartis the window end minus that. Anchoring on the window's end means an anchored or historical window reads the state as it stood then. A day clears every cadence these collectors run at by a wide margin; the slowest is database_size_stats, hourly.collection_time >= $2 AND collection_time <= $3in all three SKUs. Darling:DatabaseSizeSql,FileAutogrowthSql,DiskSpaceSql,MemoryClerkSql,PlanCacheStatsSql,MemoryStatsSqlandPgDrillDownCollector.AutogrowthPercentFilesSql, each with a byte-identical Lite twin. The Dashboard binds@lookbackStarton its six equivalents; it has no plan-cache read.MemoryStatsSqlwasn't in the issue's inventory, but it has the same shape. On Lite itsORDER BY ... LIMIT 1can't stop early across hot UNION parquet, so it sorted the whole archive.DB_CONFIGanchors on the newest capture (capture_time = MAX(capture_time)), which is the rule its drill-downConfigIssuesSqlalready used. It takes no lookback, becausedatabase_configis an on-load snapshot and its newest capture is as old as the last connect.DB_CONFIGstays present. The audit_config comments in both SKUs say so.ServerConfigSql,TraceFlagsSql,ServerMetadataSqlandServerPropertiesSqlread on-load snapshots.PlanHandleLookupSqlis an indexed per-hash lookup: 0.07 ms, 15 ms at worst.Test plan
DARLING01 before/after (SQL2022, 17 days, 18 chunks, 16 compressed;
EXPLAIN (ANALYZE, BUFFERS)viaEXECUTE, warm). Results are identical except the ghost rows:DatabaseSizeSqlMemoryClerkSqlAutogrowthPercentFilesSqlDiskSpaceSqlFileAutogrowthSqlPlanCacheStatsSqlMemoryStatsSqlThat comes to roughly 1.31 s → 66 ms per fact pass, twice per
compare_analysis. The deployed (pre-fix)get_analysis_factson SQL2022 still reportsDATABASE_TOTAL_SIZE_MB = 239067.9through MCP.Lite before/after (DuckDB 1.5.5, synthetic: 100 files, 1-minute cadence, 23 days in parquet plus 7 hot,
DatabaseSizeSql): 4,320,000 rows / 47 ms → 143,900 rows / 11 ms. The parquet arm becomesEMPTY_RESULT, because DuckDB prunes the archive on the bound parameter.New
LatestValueLookbackTests(Darling, 27). SQL-shape pins for all seven reads. A scan overPgFactCollector.AllSqlthat fails any future newest-per-series read of a cadenced collector with no lower bound, with its own positive controls. The Lite twin pinned to the same text. TheDB_CONFIGanchor.DB_CONFIG. A historical window reads the state at its own end, ghost included.New Lite
LatestValueLookbackTests(4). The same three scenarios on DuckDB, plus the Lite Recommendations never passes 24h warm-up after DuckDB archive/reset #1809 reset shape: everything in parquet, the ghost not counted, the live series archived minutes ago still counted.New Dashboard
LatestValueLookbackSqlTests(7), source-text pins, since the Dashboard has no test DB.Red-watch: with the production SQL reverted, 21/27 Darling, 4/4 Lite and 7/7 Dashboard fail. The rest are the scanner's positive controls and the constant pin.
Full
Darling.TestswithDARLING_TEST_PG(rig: PostgreSQL 18.4 + TimescaleDB 2.28.1): 12,482 total, 0 failed, 21 skipped (PGRUNTIME/SQL-gated).Full
Lite.Tests: 5,128, 0 failed. An earlier run made concurrently with the Darling suite hitStatusBarSizeReadLockTestsandAnalysisPassTokenThreadingTests("write lock never acquired", the static-lock contention family). Both pass alone and in a clean full run.Full
Dashboard.Tests: 813, 0 failed. All builds show 0 warnings.Found alongside and filed rather than folded in:
pg_server_configchunk (365-day retention). That's 7 ms at 16 chunks, but 2.5-4.9 s per read at 366 chunks on the CI-sized rig. Posture and Clock need a ruling.TRACE_FLAGSkeeps reporting a flag after it's turned off, and anchoring doesn't fix it.🤖 Generated with Claude Code
https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif