Skip to content

The chunk-interval adjustment's 1,000-chunk limit applies to each table's own chunks (#4457) - #4458

Merged
erikdarlingdata merged 4 commits into
devfrom
fix/4457-per-table-chunk-cap
Sep 27, 2026
Merged

erikdarlingdata merged 4 commits into
devfrom
fix/4457-per-table-chunk-cap

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Fixes #4457.

Why

The daily raw-hypertable chunk-interval adjustment refuses to shrink any table's chunks once the STORE-WIDE chunk count passes 1,000, no matter how far under that limit any individual table sits and no matter how far over its own memory budget the store is. A production store with 72 hypertables and 1,210 chunks in total, but a largest table of only 73 chunks, sat 2.55x over its budget and the daily adjustment moved nothing.

The design review behind this called for a per-table check instead, matching TimescaleDB's own per-hypertable guidance: a table is held back only when moving it would push its OWN forecast chunk count over the limit.

What changes

  • Each table's own current chunk count now flows into the planner as an input (timescaledb_information.hypertables.num_chunks).
  • The cap moves from a store-wide total to a per-table forecast: a table narrows one rung only while its forecast chunk count at the narrower interval stays at or under 1,000. A table whose forecast would exceed that is held, with a reason naming its forecast count, the narrower rung, and the cap.
  • The store-wide chunk total is no longer read or used as a cap input. Widening (moving to a wider interval) was never capped and stays that way.
  • Every held table's decision (not logged before) and a per-run summary (tables evaluated, tables moved, store total, budget, and the ratio between them; before, a Debug line on a run that moved nothing) now log at Information, so an operator can see why a table is or isn't moving without turning on verbose logging.

The replaced pin and why

ChunkCountCap_BlocksNarrowingEvenWhenOverBudget asserted the old store-wide behavior directly, so it was replaced rather than adapted; ChunkCountCap_NeverBlocksWidening was converted to the per-table form. New pins cover the exact per-table boundary (a table forecasting exactly 1,000 chunks narrows; one forecasting 1,001+ holds), the fact that a large store-wide total no longer matters on its own, and the production store's own shape (a 13-row scenario computed by hand and cross-checked against the planner's own rules).

Test plan

Unit pins (RawChunkIntervalPlannerTests): boundary at the 24h→12h and 12h→6h rungs, the store-wide-total-no-longer-matters case, widening still uncapped, and the production store's two-day shape. dotnet Darling.Tests.dll -class Darling.Tests.RawChunkIntervalPlannerTests — Total: 23, Failed: 0.

Live pin (RawChunkIntervalReconcilerLiveTests), through the product's own ReconcileAsync call path: 72 hypertables and 1,210 chunks total (largest 73), one compressed table whose own open bytes sit at 2.55x the budget. Asserts the rated table narrows to 12h, exactly one rung-history row is written, and the summary line is asserted by exact prefix, Information: raw chunk interval reconcile:, logging "1 moved" (plus a second arm with a huge budget logging "0 moved" and a ratio under 1, where the rated table's own held decision is also asserted by exact prefix, Information: raw chunk interval: rated_test held at 12h, and contains "moved within the last"). dotnet Darling.Tests.dll -class Darling.Tests.RawChunkIntervalReconcilerLiveTests — Total: 4, Failed: 0.

RED on dev at runtime — the new live test file, copied unchanged into a worktree of origin/dev @ 6bbc18b48, compiles there without modification (it uses only members ReconcileAsync already exposes on dev) and fails at runtime against dev's store-wide cap: Assert.Equal() Failure: Values differ / Expected: 1 / Actual: 0 — the store's 1,210-chunk total alone holds the rated table at 24h regardless of its own 2.55x-over-budget forecast. Total: 4, Failed: 1.

Logging mutations — on this branch, two REDs, each reverted after recording: (A) the summary's LogInformation changed to LogDebug fails the same production-shape fact — Assert.NotNull() Failure (the summary line no longer starts with Information:), Total: 4, Failed: 1; (B) the held line's LogInformation changed to LogDebug fails it the same way on the second arm's held-decision assertion, Total: 4, Failed: 1. Reverted; the class passes again.

Mutation — on this branch, changing the per-table forecast check from forecastChunks > PerTableChunkCountCap to forecastChunks >= 0 (every table counts as over the cap) turns both live facts RED: Total: 4, Failed: 2 (the end-to-end fact and the production-shape fact both fail to narrow). Reverted; both facts pass again.

Boundary mutation (from the earlier half of this change) — changing the per-table forecast comparison from > to >= turns the exact-1,000-forecast boundary case RED (the at-boundary table that should narrow now holds); reverted.

Also run: RawChunkIntervalRungHistoryRungTests — Total: 3, Failed: 0. DocCommentHygieneTests — Total: 77, Failed: 0. git grep -l "ChunkCountCapThreshold\|TotalChunkCountSql\|currentTotalChunkCount" returns nothing under Darling/.

Full suite: dotnet Darling.Tests.dll (no -class filter) — Total: 15316, Failed: 1828, Skipped: 130. The failures are concentrated in Viewer* WPF-facing classes (chart context menus, drill-down, settings, sidebar/status colors) and a handful of unrelated platform-shaped classes; none touch RawChunkInterval*. This macOS run cannot host WPF types, so failures there are expected (this run was not compared against one on dev); CI on Windows is the authority on whether this change introduces any new failure. Darling.Tests/Lite.Tests build clean (-p:EnableWindowsTargeting=true, 0 errors, 0 warnings).

Comment-only follow-up (9c7a65928, a test comment's wording): DocCommentHygieneTests + RawChunkIntervalPlannerTests — Total: 100, Failed: 0.

CHANGELOG entry

SECTION: Fixed
ENTRY:

The daily chunk-interval reconcile's 1,000-chunk anti-pattern cap
counted every table's chunks on the store as a single total, so a
store past 1,000 chunks in total held every table at 24 hours no
matter how far over its memory budget it was. The cap now forecasts
each table's own chunk count at the narrower interval and holds only
that table when its forecast would exceed 1,000.

- TableInput carries NumChunks; the store-wide total and its SQL are
  gone.
- Plan's cap check is per table: NumChunks * CurrentIntervalHours /
  narrower, rounded up, against PerTableChunkCountCap.
- ReconcileAsync logs every held decision and a per-run summary at
  Information.
- Planner unit tests cover the exact boundary (500 vs 501 chunks),
  the 12h->6h rung, the store-wide total no longer mattering, and a
  hand-computed production store shape (13 rows, 2.55x budget).
Seeds 72 hypertables and 1,210 chunks (largest 73) through ReconcileAsync's
own product path, with one compressed table's open bytes at 2.55x budget --
the store-wide chunk total no longer holds it at 24h, only its own forecast
decides it. Asserts the Information summary line on both a changed and an
unchanged run.
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 27, 2026 02:14
@erikdarlingdata
erikdarlingdata merged commit ab83195 into dev Sep 27, 2026
16 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the fix/4457-per-table-chunk-cap branch September 27, 2026 02:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant