Skip to content

Analysis latest-value reads look back a day instead of scanning a server's whole history, and stop counting dropped databases (#3896) - #3931

Merged
erikdarlingdata merged 1 commit into
devfrom
feature/3896-analysis-latest-value-reads
Sep 23, 2026
Merged

erikdarlingdata merged 1 commit into
devfrom
feature/3896-analysis-latest-value-reads

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Closes #3896.

Why

The analysis "latest value" reads took each series' newest row with an upper time bound only, or no bound at all. They are the database-size, percent-autogrowth, disk-space, memory-clerk, plan-cache and memory_stats facts, plus the autogrowth drill-down. The window function numbered every row the server had retained to keep a few dozen, and TimescaleDB could exclude no chunk. On Lite, where the v_ views union the parquet archive, nothing could prune a row group.

The answer was also wrong. "Latest ever" kept a dropped database's files in DATABASE_TOTAL_SIZE_MB until retention aged them out. The same files, each with an ALTER DATABASE for a database that no longer exists, stayed in the autogrowth drill-down. Erik's production note stands: on the largest store DatabaseSizeSql is about a tenth of analyze_server, and PlanRegressionSql (#3902) dominates. This fix is the correctness half plus that tenth.

The diagnosis is confirmed and extended by one read. DB_CONFIG counted the same five dropped databases on DARLING01: 12 databases per-database-latest against 7 in the newest capture.

What changes

  • The lookback. AnalysisContext.LatestValueLookback is 24 h, and LatestValueStart is the window end minus that. Anchoring on the window's end means an anchored or historical window reads the state as it stood then. A day clears every cadence these collectors run at by a wide margin; the slowest is database_size_stats, hourly.
  • Seven reads bind it, as collection_time >= $2 AND collection_time <= $3 in all three SKUs. Darling: DatabaseSizeSql, FileAutogrowthSql, DiskSpaceSql, MemoryClerkSql, PlanCacheStatsSql, MemoryStatsSql and PgDrillDownCollector.AutogrowthPercentFilesSql, each with a byte-identical Lite twin. The Dashboard binds @lookbackStart on its six equivalents; it has no plan-cache read. MemoryStatsSql wasn't in the issue's inventory, but it has the same shape. On Lite its ORDER BY ... LIMIT 1 can't stop early across hot UNION parquet, so it sorted the whole archive.
  • DB_CONFIG anchors on the newest capture (capture_time = MAX(capture_time)), which is the rule its drill-down ConfigIssuesSql already used. It takes no lookback, because database_config is an on-load snapshot and its newest capture is as old as the last connect.
  • Behavior change, which the issue asked for: a series with no sample within a day of the window's end yields no fact, where it used to yield a stale one. DB_CONFIG stays present. The audit_config comments in both SKUs say so.
  • Left unbounded on purpose. ServerConfigSql, TraceFlagsSql, ServerMetadataSql and ServerPropertiesSql read on-load snapshots. PlanHandleLookupSql is an indexed per-hash lookup: 0.07 ms, 15 ms at worst.

Test plan

  • DARLING01 before/after (SQL2022, 17 days, 18 chunks, 16 compressed; EXPLAIN (ANALYZE, BUFFERS) via EXECUTE, warm). Results are identical except the ghost rows:

    read before after
    DatabaseSizeSql 1,167 ms, 571,576 rows into the window, 11,184 buffers 56 ms, 33,750 rows, 3,189 buffers
      result 239,067.90 MB (37 files) 182,495.64 MB (27 files; the 10 extra were 5 dropped DBs)
    MemoryClerkSql 129 ms (114,068 rows) 7.4 ms
    AutogrowthPercentFilesSql 10.2 ms 0.7 ms
    DiskSpaceSql 9.0 ms 0.9 ms
    FileAutogrowthSql 8.5 ms 0.6 ms
    PlanCacheStatsSql 0.9 ms (plan 1.5 ms) 0.9 ms (plan 0.3 ms)
    MemoryStatsSql 0.09 ms (plan 1.4 ms) 0.02 ms (plan 0.3 ms)

    That comes to roughly 1.31 s → 66 ms per fact pass, twice per compare_analysis. The deployed (pre-fix) get_analysis_facts on SQL2022 still reports DATABASE_TOTAL_SIZE_MB = 239067.9 through MCP.

  • Lite before/after (DuckDB 1.5.5, synthetic: 100 files, 1-minute cadence, 23 days in parquet plus 7 hot, DatabaseSizeSql): 4,320,000 rows / 47 ms → 143,900 rows / 11 ms. The parquet arm becomes EMPTY_RESULT, because DuckDB prunes the archive on the bound parameter.

  • New LatestValueLookbackTests (Darling, 27). SQL-shape pins for all seven reads. A scan over PgFactCollector.AllSql that fails any future newest-per-series read of a cadenced collector with no lower bound, with its own positive controls. The Lite twin pinned to the same text. The DB_CONFIG anchor.

  • New Lite LatestValueLookbackTests (4). The same three scenarios on DuckDB, plus the Lite Recommendations never passes 24h warm-up after DuckDB archive/reset #1809 reset shape: everything in parquet, the ghost not counted, the live series archived minutes ago still counted.

  • New Dashboard LatestValueLookbackSqlTests (7), source-text pins, since the Dashboard has no test DB.

  • Red-watch: with the production SQL reverted, 21/27 Darling, 4/4 Lite and 7/7 Dashboard fail. The rest are the scanner's positive controls and the constant pin.

  • Full Darling.Tests with DARLING_TEST_PG (rig: PostgreSQL 18.4 + TimescaleDB 2.28.1): 12,482 total, 0 failed, 21 skipped (PGRUNTIME/SQL-gated).

  • Full Lite.Tests: 5,128, 0 failed. An earlier run made concurrently with the Darling suite hit StatusBarSizeReadLockTests and AnalysisPassTokenThreadingTests ("write lock never acquired", the static-lock contention family). Both pass alone and in a clean full run.

  • Full Dashboard.Tests: 813, 0 failed. All builds show 0 warnings.

Found alongside and filed rather than folded in:

🤖 Generated with Claude Code

https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif

…d stop counting dropped databases (#3896)

The database-size, percent-autogrowth, disk-space, memory-clerk, plan-cache
and memory_stats facts, plus the autogrowth drill-down, took each series'
newest row with only an upper time bound (or none). The window function
numbered every retained row, and "latest ever" kept dropped databases'
files in DATABASE_TOTAL_SIZE_MB, and their ALTER DATABASE in the drill-down,
until retention aged them out.

Each read now binds AnalysisContext.LatestValueStart (window end minus 24 h)
beside the window end, in Darling, Lite and the frozen Dashboard. DB_CONFIG
had the same ghost databases; being an on-load snapshot it anchors on the
newest capture instead, the rule its own drill-down already followed.

DARLING01 (SQL2022, 17 days): DatabaseSizeSql 1,167 ms -> 56 ms (571,576 ->
33,750 rows), 239,068 MB (37 files) -> 182,496 MB (27). MemoryClerkSql
129 ms -> 7 ms.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
@erikdarlingdata
erikdarlingdata merged commit c5be29e into dev Sep 23, 2026
15 of 16 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/3896-analysis-latest-value-reads branch September 23, 2026 01:23
erikdarlingdata added a commit that referenced this pull request Sep 23, 2026
…3920, #3927, #3931, #3932, #3940, #3942, #3946, #3947, #3950, #3952, #3955, #3956, #3957, #3964, #3965, #3966, #3968, #3972, #3975, #3979, #3980, #3981, #3983, #3984, #3985) (#3989)

The wave's fix PRs deliberately carried no CHANGELOG edits (parallel-agent hot-spot protocol); each agent reported its entry and this commit lands them together, byte-verified against origin/dev.


Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant