Skip to content

The startup watermark read reads only the last two days of collection_log, bound as a naive timestamp (#4469) - #4480

Merged
erikdarlingdata merged 5 commits into
devfrom
fix/4469-bounded-newest-reads
Sep 27, 2026
Merged

erikdarlingdata merged 5 commits into
devfrom
fix/4469-bounded-newest-reads

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Refs #4469. The floor removes the older-chunk part of the cost; the newest-chunk part is a follow-up.

Why

The connect-path watermark read (DarlingWorker.ReadCollectorWatermarksAsync) used to run an
unbounded MAX(collection_time) GROUP BY collector_name, which cannot answer without reading
every retained chunk, compressed ones included, because the newest instant per collector is not
known until the whole table has been read. The PostgreSQL log on the busiest measured store
showed the statement cancelled at its 10 s client deadline 12 times over about 3 days, mostly
just after a service start — the other cancels logged there were confounded by other load on
that host and are not counted here.

A field EXPLAIN on that store, run back to back (unbounded, then the bounded statement right
after), gives an honest before/after only for the read pattern, not the timing: the unbounded
run was cold — 4,775 ms, 3,166 hit + 21,483 read buffers, 13,957 ms of parallel-worker I/O read
time — while the bounded run right after it was warm (69 ms, 20,703 buffers, all shared hit),
because it read the same pages the unbounded run had just pulled into cache. Comparing those two
numbers directly overstates what the floor buys.

The real split, from the unbounded run's own EXPLAIN: 89% of its I/O read time (~12,417 ms) was
a Bitmap Heap Scan over the two newest, uncompressed chunks (~28,000 and ~19,000 rows for one
server, spread across ~12,000 and ~8,000 heap blocks) — cost the floor does not touch, because it
keeps the newest chunks. Only 11% (~1,540 ms) was spent walking the 30 older compressed chunks,
which the floor excludes outright via chunk exclusion on collection_time. That 11% grows with
retention and chunk count, so the floor is still worth having, but it is not the dominant cost on
this read, and a follow-up is needed for the newest-chunk heap-fetch cost.

The separate collector_state prune on this connect path is out of scope: measured on the same
store, its statement costs well under a millisecond.

The floor and its parameter type

ReadCollectorWatermarksSql adds AND collection_time >= $2 to the same MAX ... GROUP BY, with
$2 = now - WatermarkFloorLookback (2 days; the longest recurring collector cadence is 1 day, and a
collector whose last run is older than the floor is seeded exactly as an overdue one is). The floor parameter binds DateTime.UtcNow - WatermarkFloorLookback. Npgsql's default mapping
for a Kind=Utc DateTime is timestamptz, but collection_log.collection_time is a naive
timestamp — the product-wide contract (see BindActualPlanResolveParameters's naive-UTC bind
for a query snapshot's collection_time, and RawChunkIntervalReconciler's comments on the same
contract). Binding a timestamptz against a naive column forces a cast on one side, which can
defeat the chunk exclusion the floor exists to get.

The bind is DateTime.SpecifyKind(DateTime.UtcNow - WatermarkFloorLookback, DateTimeKind.Unspecified), sent as a plain timestamp, matching the column.

Confirmed on a live TimescaleDB rig with a 32-chunk hypertable (30 compressed) that the plan gets
chunk exclusion, with no cast on collection_time:

-- unbounded (old statement, for reference — reads every chunk):
Custom Scan (ColumnarScan) on _hyper_1_3_chunk
  ->  Index Scan using _hyper_1_3_chunk_compressed_server_id__ts_meta_v2_first_col_idx on _hyper_1_3_chunk_compressed
-- ... one such arm per chunk, 32 in all ...

-- bounded, naive-UTC bind (this fix):
Custom Scan (ColumnarScan) on _hyper_1_3_chunk
  Vectorized Filter: (collection_time >= '2026-09-25 16:31:49.014185'::timestamp without time zone)
  ->  Index Scan using _hyper_1_3_chunk_compressed_server_id__ts_meta_v2_first_col_idx on _hyper_1_3_chunk_compressed
        Index Cond: ((server_id = 1) AND (_ts_meta_v2_first_collection_time >= '2026-09-25 16:31:49.014185'::timestamp without time zone))

The filter and index condition compare collection_time/_ts_meta_v2_first_collection_time
(both naive timestamp columns) against a bare, uncast timestamp literal — no
::timestamptz cast anywhere in the plan — and only the newest chunks execute (DarlingWatermarkFloorPlanShapeLiveTests asserts the touched-chunk bound on the plan).

The runtime pin

DarlingWatermarkFloorScanBoundLiveTests is a live test driven ONLY through the
existing product method DarlingWorker.ReadCollectorWatermarksAsync(postgres, serverId, logger, ct) — no reference to the new ReadCollectorWatermarksSql or WatermarkFloorLookback members,
so it compiles unmodified against a build that predates both.

It seeds 30 days of collection_log history (most chunks old enough to compress, the field
shape), snapshots every chunk relation's pg_stat_user_tables scan counters — for a compressed
chunk this means the relation named by _timescaledb_catalog.compression_settings.compress_relid,
since that is where TimescaleDB's actual scan against compressed data lands — calls the method,
and asserts (a) the returned watermarks equal the unbounded oracle, same as the existing
plan-shape test, and (b) not one chunk older than the floor (minus a chunk width of slack) shows
a new scan afterward.

RED against origin/dev (this test file copied unmodified into a detached worktree,
building and running there unchanged): the pin fails at RUNTIME, not at compile time. The old
method has no floor, so an old chunk's (seq_scan, idx_scan) counters move (one new index scan):

Darling.Tests.DarlingWatermarkFloorScanBoundLiveTests.ReadCollectorWatermarksAsync_NeverScansAChunkOlderThanTheFloor [FAIL]
  Assert.Equal() Failure: Values differ
  Expected: Tuple (46, 34)
  Actual:   Tuple (46, 35)
Darling.Tests  Total: 1, Errors: 0, Failed: 1, Skipped: 0, Not Run: 0, Time: 0.608s

GREEN on this branch:

Darling.Tests  Total: 1, Errors: 0, Failed: 0, Skipped: 0, Not Run: 0, Time: 0.625s

The mutation (removed the AND collection_time >= $2 predicate from
ReadCollectorWatermarksSql, rebuilt, reran):

Darling.Tests.DarlingWatermarkFloorScanBoundLiveTests.ReadCollectorWatermarksAsync_NeverScansAChunkOlderThanTheFloor [FAIL]
  Assert.Equal() Failure: Values differ
  Expected: Tuple (54, 41)
  Actual:   Tuple (54, 42)
Darling.Tests  Total: 1, Errors: 0, Failed: 1, Skipped: 0, Not Run: 0, Time: 0.651s

Reverted immediately after; the diff in this PR does not carry the mutation.

Test plan

  • DarlingWatermarkFloorPlanShapeLiveTests, DarlingWatermarkSeedLiveTests,
    DarlingWatermarkFloorScanBoundLiveTests (new), DocCommentHygieneTests: 80/80 passing
    together on a live TimescaleDB 2.30.1 rig, 0 warnings on the build.
  • RED confirmed at runtime against origin/dev (above).
  • Mutation confirmed RED, then reverted (above).
  • After the final doc-comment commit (no code change): DocCommentHygieneTests 77/77; Darling.Tests
    and Lite.Tests build with 0 warnings.

CHANGELOG

SECTION: Changed
ENTRY:

…nded MAX/GROUP BY (#4469)

The connect-path read of every collector's last run touched every retained collection_log chunk, compressed ones included, before it could find the newest instant per collector. On the busiest measured store that took 4,775 ms and read 21,483 buffers against a 32-chunk hypertable; a literal 2-day floor on collection_time drops that to 69 ms with the same statement shape, because TimescaleDB can now exclude every older chunk before opening it.

The floor is safe because this read only seeds each collector's next-due time from its own cadence: a collector whose true last run predates the floor is, by definition, already overdue on every schedule this product ships (the longest recurring cadence is 1 day), so it falls back to the same never-run seed the read already used on an outright failure.
#4469)

- The floor parameter now binds via DateTime.SpecifyKind(..., Unspecified)
  instead of the default Kind=Utc mapping, matching the naive timestamp
  column collection_time and the same product-wide contract other binds use
  (BindActualPlanResolveParameters). Confirmed on a live TimescaleDB rig: the
  bounded plan still shows chunk exclusion on the recent chunks, with the
  comparison as timestamp >= timestamp (no cast on collection_time).
- Added DarlingWatermarkFloorScanBoundLiveTests, a runtime pin driven only
  through the existing product method ReadCollectorWatermarksAsync (no new
  member reference), which snapshots pg_stat_user_tables scan counters
  (including each compressed chunk's compression_settings.compress_relid
  relation) before and after the call and asserts no chunk older than the
  floor shows a new scan, plus that the returned watermarks equal the
  unbounded oracle. This test fails at RUNTIME against the pre-fix method,
  which cannot avoid scanning every chunk.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant