Skip to content

The start-up hole scan probes each bucket once instead of joining the whole materialization and the whole source table (#3933) - #3972

Merged
erikdarlingdata merged 1 commit into
devfrom
feature/3933-hole-scan-probes
Sep 23, 2026
Merged

erikdarlingdata merged 1 commit into
devfrom
feature/3933-hole-scan-probes

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Closes #3933.

Why

TimescaleSupport.MaterializationHoleScanSql is the scan behind the start-up materialization-hole repair (#3731). RepairMaterializationHolesAsync runs it for every one of the 23 registered aggregates on every service start, and again after each repair.

Its doc said the two correlated probes are "each an index range ... so cost follows the number of buckets scanned and not the size of either table". The planner did something else. Written as a bare NOT EXISTS and EXISTS, both are pulled up into joins, and neither join can push the per-bucket bound into its scan:

Confirmed on DARLING01, a real store with a few outage hours in its 35-day baseline span. The wait-stats baseline's scan materialized all 9,785,051 wait_stats rows to disk (21,501 temp blocks written, 64,503 read) and read them four times: 4.5 s for one aggregate. The whole start-up scan across 23 aggregates was 5.9 s. On a 43-server store with 30 days of raw data, the same shape is several times the rows on every start.

What changes

  • The fence: each probe carries OFFSET 0. PostgreSQL's simplify_EXISTS_query refuses to pull a sub-select with an OFFSET up into a join ("OFFSET 0 ... traditionally is used as an optimization fence"). So each probe runs as a SubPlan, once per bucket, with the chunk it needs picked at run time.
  • Probe order: the candidate buckets are fenced the same way, so the source is probed only for the buckets the materialization probe found empty.
  • Results are unchanged: the same buckets in the same order, checked against the old text as an oracle, below.
  • Docs: the method doc records the finding, and the type summary's "both are index probes" now points at the fence that makes it true.
  • Scope: no store migration, and no change to the pass's logic, cap, logging or tally.

Test plan

  • DARLING01, read-only as -U admin with default_transaction_read_only = on. Every one of the 23 targets over the span the pass itself would scan, old text vs new, EXPLAIN (ANALYZE, BUFFERS, TIMING OFF), median of 2. Same buckets on every target.

    aggregate before after
    wait_stats_interval_baseline 4,662 ms (64,503 temp blocks read) 22.5 ms
    perfmon_interval_baseline 999 ms (401,386 buffers) 23.8 ms
    memory_baseline 76.8 ms 8.4 ms
    deadlock_baseline 61.2 ms 4.3 ms
    session_stats_baseline 41.7 ms 6.2 ms
    query_store_stats_interval_daily 14.2 ms 0.1 ms
    the other 17 ≤ 6 ms each ≤ 1.4 ms each
    all 23: one start's scans 5,876 ms 75 ms
  • Rig, one busy server (571K raw query_stats rows, 807K hourly rollup rows) with a six-hour outage and a two-hour tail planted: query_stats_hourly's scan went from 3,886 ms to 6.3 ms, finding the same two buckets.

  • New tests:

  • Red on the old shape: with the product file reverted, the fence pin and the plan test fail (the old plan's nested-loop joins). The oracle test passes trivially then, as it must.

  • Existing pins updated deliberately (text, not intent): MaterializationHoleRepairTests.TheSourceFilter_IsTheAggregatesOwnWhere_Verbatim_OrEmpty. Both repair live tests (tail, outage, forced-refresh hole, successor filter, cap and deferral tally) pass on the new text.

  • Full Darling.Tests against the rig with DARLING_TEST_PG: 12,703 total, 0 failed, 24 skipped (env-gated managed-runtime, PostgreSQL-target, live SQL Server and symlink tests), on dev db93e292; after rebasing onto 8395bece (A restart no longer raises a false Store Checkpointer Pressure warning or reads as a WAL-forced checkpoint on a monitored PostgreSQL server (#3955) #3969, unrelated), the hole-repair, store-convergence, not-carried and doc-hygiene classes were re-run (110 tests, 0 failed)

🤖 Generated with Claude Code

https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif

… whole materialization and the whole source (#3933)

The materialization-hole scan's two per-bucket EXISTS probes were pulled up
into joins: an anti-join that read the whole materialization hypertable,
and a semi-join over a Materialize of the whole source table, read once per
bucket with no materialized row. Every outage hour in the scan span is such
a bucket. On DARLING01 the wait-stats baseline's scan materialized all
9,785,051 wait_stats rows to disk and read them four times: 4.5 s for one
aggregate. Each probe now carries an OFFSET 0 fence and runs as a per-bucket
SubPlan, and the source is probed only for buckets the materialization probe
found empty. The found buckets are identical: a live test runs the old text
as an oracle on four kinds of target.

DARLING01, all 23 aggregates, one start's scans: 5,876 ms -> 75 ms.
Rig, one busy server after an outage: 3,886 ms -> 6.3 ms for the hourly.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
@erikdarlingdata
erikdarlingdata enabled auto-merge (squash) September 23, 2026 02:05
@erikdarlingdata
erikdarlingdata merged commit c67eb7e into dev Sep 23, 2026
25 of 28 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/3933-hole-scan-probes branch September 23, 2026 02:58
erikdarlingdata added a commit that referenced this pull request Sep 23, 2026
…ection-health policy run waits out a launch that beat the park (#3982)

* Chunk-exclusion tests count a bitmap plan's chunks once, and the collection-health policy run waits out a launch that beat the park

Two CI flakes that blocked #3968 and #3972 on code neither PR touched.

LatestValueLookbackLivePostgresTests reported "planned 4 chunks" for a
24-hour read whose plan scanned two. TimescaleDB names a chunk's indexes
after the chunk, so under a bitmap plan the "Bitmap Index Scan on
_hyper_12_11_chunk_..._idx" child matched the chunk pattern too, and the
count doubled whenever the planner picked bitmap scans. The same pattern
lived in FleetReadsAreBoundedByTheFleetTests and CaptureDownChunkOrderTests.
One shared PlanChunkScans helper now ends the match on the chunk name, and
PlanChunkScansTests pins it against the plan CI printed.

CollectionHealthAggregateTests' run_job failed 55P03 "concurrent refresh":
collection_health_hourly's policy has no initial_start, so TimescaleDB
launches its first run the moment the product commits it, before the
fixture can park it. RunPolicyAsync now retries 55P03 for up to 30 s.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

* The chunk counters #3980 added go through PlanChunkScans too

#3980 merged to dev with two more ChunkScans callers in
LatestValueLookbackTests and a fourth copy of the double-counting pattern
in PgTargetConfigSnapshotBoundTests. Both now use the shared helper.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
erikdarlingdata added a commit that referenced this pull request Sep 23, 2026
…3920, #3927, #3931, #3932, #3940, #3942, #3946, #3947, #3950, #3952, #3955, #3956, #3957, #3964, #3965, #3966, #3968, #3972, #3975, #3979, #3980, #3981, #3983, #3984, #3985) (#3989)

The wave's fix PRs deliberately carried no CHANGELOG edits (parallel-agent hot-spot protocol); each agent reported its entry and this commit lands them together, byte-verified against origin/dev.


Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant