Skip to content

Derive each raw hypertable's chunk interval from ingest (#4211) - #4332

Merged
erikdarlingdata merged 3 commits into
devfrom
feature/4211-chunk-interval-derivation
Sep 25, 2026
Merged

erikdarlingdata merged 3 commits into
devfrom
feature/4211-chunk-interval-derivation

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Part of #4211.

Why

The #4211 ruling (issuecomment-5836205190) derives each raw hypertable's chunk_time_interval (I) from
ingest instead of leaving every table at a fixed 1-day width, while keeping compress_after (A) fixed at
1 day everywhere. This lane (I1) builds the pure derivation only: no database reads, no
set_chunk_time_interval calls, no reconcile loop. Lane I2 wires it into the TimescaleDB ensure path.

What changes

  • New RawChunkIntervalPlanner (Darling/PerformanceMonitor.Darling.Storage/RawChunkIntervalPlanner.cs),
    a pure static class. Plan(tables, budgetBytes, currentTotalChunkCount, asOfUtc) takes plain values (name,
    ingest bytes/hour, current interval, when it last changed) for every raw hypertable and returns one
    Decision per table (target interval + a one-line reason). Rules, all from the ruling:
    • Ladder: 24h / 12h / 6h only. The ceiling is read from TimescaleSupport.ChunkIntervalDays * 24, not
      restated as a literal.
    • The budget covers every table's open chunk (rate × its own interval, summed), not just the heavy ones
      (ruling decision 2/review H3).
    • Which tables narrow: heaviest ingest rate first, one rung per table per call, until the total fits, every
      eligible table is at the floor, or the chunk-count cap blocks the next move.
    • A table changed within the last day is held in place, in either direction (one-rung-per-day, ruling
      decision 4).
    • Widening (hysteresis): only for a table the narrowing pass left alone, one rung, and only when the
      store-wide open-chunk total WITH that move made stays at or under half of B. Lightest ingest rate first;
      a rate tie breaks on the table name (ruling issuecomment-5836734285).
    • Chunk-count cap: ChunkCountCapThreshold = 1_000. Since the planner has no per-table retention input, it
      is a coarse store-wide brake (refuses every narrowing move once the store's current total chunk count is
      at or past the threshold), not a per-table forecast. It never blocks widening.
    • A current interval outside {24,12,6} (a hand-set legacy value on an adopted store) is left unchanged in
      both directions rather than guessed at.
    • ManagedBudgetBytes(totalPhysicalMemoryBytes) = 25% of the SAME raw RAM figure
      DarlingManagedPostgres.DeriveMemorySettings receives — explicitly NOT that method's SharedBuffersMb
      output, which is capped at 1 GB for the co-located-store/Windows 487 mitigation (Managed store: cap shared_buffers at 1GB for the co-located reality (v5) + MaxPoolSize=24 — heals the Windows 487 spawn failures #1559).
    • BringYourOwnBudgetBytes(sharedBuffersBytes, effectiveCacheSizeBytes) = max(shared_buffers, effective_cache_size / 3), per the review's H3 fix (plain shared_buffers alone would make almost every
      table "heavy" on an untuned 128 MB default).
  • EventWindowFloor.SkewAllowance now its own TimeSpan.FromDays(1) constant instead of reading
    TimescaleSupport.ChunkIntervalDays. Same value today; decoupled because it's a clock-skew tolerance, not a
    chunk-width reference, and the three tables it guards (blocked_process_reports, dmv_blocking_snapshots,
    deadlocks) are not among the tables this planner narrows, so nothing changes in effect.
  • Pins updated on purpose:
    • TimescaleSupport.ChunkIntervalDays and CompressAfterDays doc comments reworded: they are now the
      ladder's ceiling, not "the" uniform width/delay. Listed dependents that stay correct unmodified because
      they already treat ChunkIntervalDays as a margin/ceiling: DarlingDimensionGcBound,
      QueryStorePlanMap.MarginOrderingHolds, RollupBackfill.SliceWidth, the aggregate compress margin.
    • RollupBackfill.SliceWidth's "Compression waits for the next fixed tick, so a just-closed chunk can sit uncompressed most of a day (81 GB observed); deadlock correlation to watch #1778 slice note" reworded: a slice is no longer guaranteed to be exactly
      one chunk of a narrowed table, but stays bounded to at most one day, which is what the lock-safety
      argument actually needs.
    • DarlingRetentionTests and TimescaleSupportTests comments near the ChunkIntervalDays/CompressAfterDays
      pins reworded to say why they still hold (values unchanged; only the meaning of "the ladder's ceiling"
      is new).
    • New pin: CompressAfterDays_IsAtLeastTheHourlyRefreshStartOffset in TimescaleSupportTests.cs asserts
      TimeSpan.FromDays(CompressAfterDays) >= HourlyRefreshStartSpan, comparing spans rather than re-parsing
      the two "1 day" strings.

Ruled (issuecomment-5836734285): a table widens one rung only when the store-wide open-chunk total, with that move made, stays at or under half of B — not the equal-share definition this PR shipped with originally.

Test plan

  • Darling/Darling.Tests/Darling.Tests.csproj builds with 0 Warning(s), 0 Error(s).
  • New RawChunkIntervalPlannerTests (15 tests): small store/nothing moves, Store A's design numbers
    (R = 2.2 GB/h, 1.62 GB/h in three query tables) on a 63 GiB host (3 of 4 tables narrow to 6h, the
    lightest holds at 12h under the hysteresis check) and on a 31.5 GiB host (all four reach the 6h floor
    and the store still doesn't fit — the honest outcome review finding H3 predicted for a host this size),
    one-rung-per-day (including the exact-24-hours-ago boundary), hysteresis on the way up (including the
    never-past-ceiling case), the chunk-count cap (blocks narrowing, never blocks widening), an
    outside-the-ladder current interval, the budget helpers, and the "every table counts, not just the
    heavy ones" invariant. All computed by hand against the algorithm before running, then verified.
  • RawChunkIntervalPlannerTests, TimescaleSupportTests, DarlingRetentionTests,
    DocCommentHygieneTests, FleetReadsAreBoundedByTheFleetTests (the class asserting SkewAllowance):
    214 total, 0 failed, 19 skipped (live-database tests, no rig started per this lane's brief).
  • Merged origin/dev (9cb990a) after the initial commit; no conflicts, rebuilt and re-ran the same
    classes clean.
  • Full suite: not run here — CI runs it.
  • Live/rig tests: none apply; this lane is pure code with no database wiring.

What lane I2 must wire

  • Read each raw hypertable's rate from chunk_compression_stats() (not the 2.29-only catalog columns —
    review finding M1), the store's current chunk_time_interval per table, and when it last moved.
  • Call RawChunkIntervalPlanner.Plan with those inputs plus ManagedBudgetBytes/BringYourOwnBudgetBytes and
    the store's current total chunk count.
  • For each Decision where Changes is true, call set_chunk_time_interval only where the catalog differs
    (as the existing materialization converge already does), and persist the rung-change timestamp somewhere
    durable so IntervalLastChangedUtc survives a restart (ruling/review finding M4 — a restart loop must not
    be able to move a rung twice).
  • The daily reconcile, its off switch, its rung-history log line, and the UTC time-zone guard (ruling decisions
    6 and the review's M4) — none of that exists yet; this lane only produces the per-call decision.
  • compress_after itself is NOT part of this wiring: it stays fixed at CompressAfterDays on every table.
    Nothing here calls alter_job for it.

What the coordinator should double-check

  • The widening rule matches the ruling (issuecomment-5836734285): store total after the move <= B/2, lightest first.
  • That I2's wiring reads IngestBytesPerHour as heap-plus-index bytes (review finding H3), not a heap-only
    figure, and sources it from chunk_compression_stats() per review finding M1.
  • No CHANGELOG entry: no user-visible change ships until I2 wires this in.

🤖 Generated with Claude Code

https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3

erikdarlingdata and others added 2 commits September 25, 2026 13:31
Adds RawChunkIntervalPlanner, a pure static planner (#4211): given each raw
hypertable's ingest rate, current chunk_time_interval and a RAM-derived
budget, it decides a target interval on a 24h/12h/6h ladder, one rung per
table per day, budget-aware across every table (not only the "heavy" ones),
with hysteresis on the way back up and a store-wide chunk-count cap. No
database wiring yet - wiring this into the TimescaleDB ensure path and
compress_after is a separate change.

Per the #4211 ruling (issuecomment-5836205190): compress_after stays fixed
at 1 day everywhere regardless of a table's derived interval; only
chunk_time_interval moves. ChunkIntervalDays and CompressAfterDays become
the ladder's ceiling rather than "the" uniform chunk width, so their doc
comments, a retention test comment and RollupBackfill's #1778 slice note are
reworded to match. EventWindowFloor.SkewAllowance gets its own 1-day
constant instead of reading ChunkIntervalDays. A new pin asserts
compress_after stays at least HourlyRefreshStartOffset.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
The #4211 ruling (issuecomment-5836734285) replaces the per-table equal
share (budgetBytes / tables.Count) with a store-wide check: a table
moves back up one rung only when the store-wide open-chunk total, with
that move made, stays at or under half of B. With about 72 hypertables
the equal share almost never let a narrowed heavy table widen again.
Widening now considers the lightest-rate table first, same tie-break
as narrowing.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 25, 2026 17:52
@erikdarlingdata
erikdarlingdata enabled auto-merge (squash) September 25, 2026 17:52
@erikdarlingdata
erikdarlingdata merged commit 35e2d06 into dev Sep 25, 2026
18 of 20 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/4211-chunk-interval-derivation branch September 25, 2026 18:08
erikdarlingdata added a commit that referenced this pull request Sep 26, 2026
…4438)

Adds the missing [Unreleased] CHANGELOG entries for nine merged PRs. Two more need none.

- Fixed: #3588, #3886, #3889, #3900, #3911, #4016, #4032 and #4042.
- Changed: #4157. llms.txt and CITATION.cff now match the shipped product.
- None:
  - #4332 adds RawChunkIntervalPlanner without wiring it in; the later PR that wires it carries its own entry.
  - #4398 changes tests only.
- Each entry was written from its PR's diff, with a [#N] label linking the pull request.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant