Repository navigation
Derive each raw hypertable's chunk interval from ingest (#4211) - #4332
Merged
Merged
Conversation
Adds RawChunkIntervalPlanner, a pure static planner (#4211): given each raw hypertable's ingest rate, current chunk_time_interval and a RAM-derived budget, it decides a target interval on a 24h/12h/6h ladder, one rung per table per day, budget-aware across every table (not only the "heavy" ones), with hysteresis on the way back up and a store-wide chunk-count cap. No database wiring yet - wiring this into the TimescaleDB ensure path and compress_after is a separate change. Per the #4211 ruling (issuecomment-5836205190): compress_after stays fixed at 1 day everywhere regardless of a table's derived interval; only chunk_time_interval moves. ChunkIntervalDays and CompressAfterDays become the ladder's ceiling rather than "the" uniform chunk width, so their doc comments, a retention test comment and RollupBackfill's #1778 slice note are reworded to match. EventWindowFloor.SkewAllowance gets its own 1-day constant instead of reading ChunkIntervalDays. A new pin asserts compress_after stays at least HourlyRefreshStartOffset. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
The #4211 ruling (issuecomment-5836734285) replaces the per-table equal share (budgetBytes / tables.Count) with a store-wide check: a table moves back up one rung only when the store-wide open-chunk total, with that move made, stays at or under half of B. With about 72 hypertables the equal share almost never let a narrowed heavy table widen again. Widening now considers the lightest-rate table first, same tie-break as narrowing. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
erikdarlingdata
marked this pull request as ready for review
September 25, 2026 17:52
erikdarlingdata
enabled auto-merge (squash)
September 25, 2026 17:52
This was referenced Sep 25, 2026
erikdarlingdata
added a commit
that referenced
this pull request
Sep 26, 2026
…4438) Adds the missing [Unreleased] CHANGELOG entries for nine merged PRs. Two more need none. - Fixed: #3588, #3886, #3889, #3900, #3911, #4016, #4032 and #4042. - Changed: #4157. llms.txt and CITATION.cff now match the shipped product. - None: - #4332 adds RawChunkIntervalPlanner without wiring it in; the later PR that wires it carries its own entry. - #4398 changes tests only. - Each entry was written from its PR's diff, with a [#N] label linking the pull request.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of #4211.
Why
The #4211 ruling (issuecomment-5836205190) derives each raw hypertable's
chunk_time_interval(I) fromingest instead of leaving every table at a fixed 1-day width, while keeping
compress_after(A) fixed at1 day everywhere. This lane (I1) builds the pure derivation only: no database reads, no
set_chunk_time_intervalcalls, no reconcile loop. Lane I2 wires it into the TimescaleDB ensure path.What changes
RawChunkIntervalPlanner(Darling/PerformanceMonitor.Darling.Storage/RawChunkIntervalPlanner.cs),a pure static class.
Plan(tables, budgetBytes, currentTotalChunkCount, asOfUtc)takes plain values (name,ingest bytes/hour, current interval, when it last changed) for every raw hypertable and returns one
Decisionper table (target interval + a one-line reason). Rules, all from the ruling:TimescaleSupport.ChunkIntervalDays * 24, notrestated as a literal.
(ruling decision 2/review H3).
eligible table is at the floor, or the chunk-count cap blocks the next move.
decision 4).
store-wide open-chunk total WITH that move made stays at or under half of B. Lightest ingest rate first;
a rate tie breaks on the table name (ruling issuecomment-5836734285).
ChunkCountCapThreshold = 1_000. Since the planner has no per-table retention input, itis a coarse store-wide brake (refuses every narrowing move once the store's current total chunk count is
at or past the threshold), not a per-table forecast. It never blocks widening.
both directions rather than guessed at.
ManagedBudgetBytes(totalPhysicalMemoryBytes)= 25% of the SAME raw RAM figureDarlingManagedPostgres.DeriveMemorySettingsreceives — explicitly NOT that method'sSharedBuffersMboutput, which is capped at 1 GB for the co-located-store/Windows 487 mitigation (Managed store: cap shared_buffers at 1GB for the co-located reality (v5) + MaxPoolSize=24 — heals the Windows 487 spawn failures #1559).
BringYourOwnBudgetBytes(sharedBuffersBytes, effectiveCacheSizeBytes)=max(shared_buffers, effective_cache_size / 3), per the review's H3 fix (plainshared_buffersalone would make almost everytable "heavy" on an untuned 128 MB default).
EventWindowFloor.SkewAllowancenow its ownTimeSpan.FromDays(1)constant instead of readingTimescaleSupport.ChunkIntervalDays. Same value today; decoupled because it's a clock-skew tolerance, not achunk-width reference, and the three tables it guards (
blocked_process_reports,dmv_blocking_snapshots,deadlocks) are not among the tables this planner narrows, so nothing changes in effect.TimescaleSupport.ChunkIntervalDaysandCompressAfterDaysdoc comments reworded: they are now theladder's ceiling, not "the" uniform width/delay. Listed dependents that stay correct unmodified because
they already treat
ChunkIntervalDaysas a margin/ceiling:DarlingDimensionGcBound,QueryStorePlanMap.MarginOrderingHolds,RollupBackfill.SliceWidth, the aggregate compress margin.RollupBackfill.SliceWidth's "Compression waits for the next fixed tick, so a just-closed chunk can sit uncompressed most of a day (81 GB observed); deadlock correlation to watch #1778 slice note" reworded: a slice is no longer guaranteed to be exactlyone chunk of a narrowed table, but stays bounded to at most one day, which is what the lock-safety
argument actually needs.
DarlingRetentionTestsandTimescaleSupportTestscomments near theChunkIntervalDays/CompressAfterDayspins reworded to say why they still hold (values unchanged; only the meaning of "the ladder's ceiling"
is new).
CompressAfterDays_IsAtLeastTheHourlyRefreshStartOffsetinTimescaleSupportTests.csassertsTimeSpan.FromDays(CompressAfterDays) >= HourlyRefreshStartSpan, comparing spans rather than re-parsingthe two "1 day" strings.
Ruled (issuecomment-5836734285): a table widens one rung only when the store-wide open-chunk total, with that move made, stays at or under half of B — not the equal-share definition this PR shipped with originally.
Test plan
Darling/Darling.Tests/Darling.Tests.csprojbuilds with 0 Warning(s), 0 Error(s).RawChunkIntervalPlannerTests(15 tests): small store/nothing moves, Store A's design numbers(R = 2.2 GB/h, 1.62 GB/h in three query tables) on a 63 GiB host (3 of 4 tables narrow to 6h, the
lightest holds at 12h under the hysteresis check) and on a 31.5 GiB host (all four reach the 6h floor
and the store still doesn't fit — the honest outcome review finding H3 predicted for a host this size),
one-rung-per-day (including the exact-24-hours-ago boundary), hysteresis on the way up (including the
never-past-ceiling case), the chunk-count cap (blocks narrowing, never blocks widening), an
outside-the-ladder current interval, the budget helpers, and the "every table counts, not just the
heavy ones" invariant. All computed by hand against the algorithm before running, then verified.
RawChunkIntervalPlannerTests,TimescaleSupportTests,DarlingRetentionTests,DocCommentHygieneTests,FleetReadsAreBoundedByTheFleetTests(the class assertingSkewAllowance):214 total, 0 failed, 19 skipped (live-database tests, no rig started per this lane's brief).
origin/dev(9cb990a) after the initial commit; no conflicts, rebuilt and re-ran the sameclasses clean.
What lane I2 must wire
chunk_compression_stats()(not the 2.29-only catalog columns —review finding M1), the store's current
chunk_time_intervalper table, and when it last moved.RawChunkIntervalPlanner.Planwith those inputs plusManagedBudgetBytes/BringYourOwnBudgetBytesandthe store's current total chunk count.
DecisionwhereChangesis true, callset_chunk_time_intervalonly where the catalog differs(as the existing materialization converge already does), and persist the rung-change timestamp somewhere
durable so
IntervalLastChangedUtcsurvives a restart (ruling/review finding M4 — a restart loop must notbe able to move a rung twice).
6 and the review's M4) — none of that exists yet; this lane only produces the per-call decision.
compress_afteritself is NOT part of this wiring: it stays fixed atCompressAfterDayson every table.Nothing here calls
alter_jobfor it.What the coordinator should double-check
IngestBytesPerHouras heap-plus-index bytes (review finding H3), not a heap-onlyfigure, and sources it from
chunk_compression_stats()per review finding M1.🤖 Generated with Claude Code
https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3