Skip to content

Store rung V152: drop the Query Store rollups' unread group indexes - #4506

Merged
erikdarlingdata merged 9 commits into
devfrom
feat/v152-drop-unread-cagg-group-indexes
Sep 28, 2026
Merged

erikdarlingdata merged 9 commits into
devfrom
feat/v152-drop-unread-cagg-group-indexes

Conversation

@erikdarlingdata

@erikdarlingdata erikdarlingdata commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

Refs #4503.

Why

On one production store, the Query Store rollups' refreshes wrote ~30% of the store's WAL over ~71 hours, and most of it went into indexes nothing reads.

  • The indexes: TimescaleDB creates a "group" index (column, bucket DESC) on a continuous aggregate's materialization hypertable for every GROUP BY column, unless the aggregate says timescaledb.create_group_indexes = false.
    • query_store_stats_hourly, _corrected_hourly, _daily, _corrected_daily and _daygrain_daily each carry 5 (database_name, module_name, query_hash, server_id, server_name).
    • query_store_stats_interval_daily carries 11 (~10.9 GB). It wrote 34 GB of WAL in 3 refreshes.
  • Their reads: idx_scan summed over each index's chunks, over the store's lifetime, is 0 on the real chunks for every group index except server_id (and server_name on one hourly rollup: 131).
    • One empty 8 KB chunk on each hourly rollup shows scans on its database_name index (1,498 and 1,976), and the same pattern on server_id (5,397 and 2,848). On an empty chunk every index costs the same, so which one the planner picks there says nothing. The likely source is the trend read's optional database filter.
  • The precedent: Alert reads die during the interval-hourly Query Store refresh: 26/26 runs over four minutes, 417 s average, breathing with load — the grid protects writers from locks, nothing protects readers from I/O #3597 measured 4.3× the WAL from these indexes on query_store_stats_interval_hourly, and turned them off there only.

What changes

  • Store rung V152 drops the unread group indexes on the six rollups. It keeps each one's bucket index and (server_id, bucket DESC), plus (server_name, bucket DESC) on the two hourly rollups.
    • Matched by key columns through the catalog (pg_index / pg_attribute on the materialization hypertable resolved from the view), never by a built name. PostgreSQL truncates identifiers at 63 bytes, so …_runtime_stats_interval_id_bucket_idx doesn't exist under that name, and a name-based DROP INDEX IF EXISTS would silently keep it.
    • It's idempotent, and skips a rollup that doesn't exist.
  • Fresh stores match upgraded ones exactly. The six rollups are created with timescaledb.create_group_indexes = false, then the kept indexes are created unnamed ON the rollup as (column, bucket DESC). That gives PostgreSQL's default name, which is TimescaleDB's own form, so a fresh store's indexes are byte-identical to an upgraded store's.
  • The readers and the index each one uses after V152 (every non-test reader of the six rollups):
reader rollup filter index after V152
QueryStoreTrendRouting.BuildRollupTrendSql (Viewer and MCP trend reads) _corrected_hourly server_id, bucket range, optional database_name (server_id, bucket)
RollupBoundsSql _corrected_hourly min/max(bucket) bucket
Composed panels (ComposeCompiler via ComposeSourceRouter) the hourly and daily rollups and _daygrain_daily bucket range; server_name only when the panel is scoped to servers; optional database, module or query hash filters hourly: (server_name, bucket) or bucket; daily: bucket
Backfill and materialization-hole reads (RollupBackfill, MaterializationHoles) all six min/max(bucket) bucket
none (it only feeds _daygrain_daily) _interval_daily (server_id, bucket) kept
  • No reader filters a dropped column without a bucket range. The daily rollups' server_name index is dropped too: it showed 0 lifetime scans, so a scoped daily panel already runs on bucket plus a filter today.
  • A composed panel that filters on a dropped column still returns the same rows, through the kept indexes.
  • The Viewer's store probe recognises V152 by the absence of the hourly rollup's (query_hash, bucket) index AND V151's own marker. It walks the view's own dependency (pg_rewrite / pg_depend), so it needs no extension check, and it works as the viewer role. Requiring V151's marker keeps a store below V151 without the rollup from reading as V152. A plain-PostgreSQL store at V151 reads as V152, which is harmless: V152 does nothing without TimescaleDB, so the two schemas are identical.
  • Without TimescaleDB, V152 does nothing and still records version 152 (there are no rollups to change).
  • Locks: DROP INDEX on a hypertable locks it and every chunk, inside the rung's one transaction. V152 sets lock_timeout to the migration statement timeout minus 20 seconds (280 s today, derived from the same constant so the two can't cross), long enough to wait out a running hourly rollup refresh (about 264 s). During a one-time upgrade start that lands on a running rollup refresh, Query Store trend reads can wait up to ~4.5 minutes, and collection starts after the upgrade finishes. The 120-second start retry budget only decides whether to start another attempt; it never cancels one in flight.
  • The drop skips UNIQUE indexes and indexes with INCLUDE columns, so an index an operator added on the same columns is left alone.
  • One test guard is relaxed: MigrationDataMovingRungCensusPins.TheLadderStillContainsNoDynamicSql now exempts V152 (and only V152, through a named list with its reason). V152's DROP INDEX has to be built at runtime, because the index names differ per store. The census exists to account for data-moving DDL against the migration lock budget, and a DROP INDEX moves no data.
  • Lite has no continuous aggregates, so there's no Lite twin.

Test plan

In-process on macOS against a live PostgreSQL 18 + TimescaleDB 2.30.

  • CaggGroupIndexDropLiveTests:
    • (a) after a fresh migrate, each of the six rollups has exactly the kept indexes (matched by key columns) and none of the dropped ones. RED at runtime on dev.
    • (c) V152 run twice leaves the same pg_indexes rows.
  • CaggGroupIndexDropUpgradeLiveTests: a store at V151 with the six rollups materialized and compressed.
    • Every reader query's rows are captured, V152 runs, and the rows are captured again. The rows are identical (ordered), and the index assertions hold.
    • The readers: BuildRollupTrendSql with and without the database filter, RollupBoundsSql, and a server + bucket + database read per rollup.
    • RED at runtime on dev.
  • A mutation: matching the drop by a built name instead of key columns → RED; the runtime_stats_interval_id index survives because of the 63-byte truncation.
  • PgSchemaGeneratorTests: fresh vs upgraded, extended to these indexes.
  • The rung: CaggGroupIndexDropRungTests (the new top rung: ladder, probe sentinel and arm), and AgGroupIdRungTests in the stays-true shape. The "zero every rung above me" loops pick up the new ordinal unchanged.
  • Three pins widened for the new design: two TimescaleContinuousAggregateTests assertions expect the new WITH (…) clause, and IntervalDedupMaterializationIndexesTests now names the six rollups V152 measures. It still asserts every other aggregate keeps its default group indexes.
  • Added since: V152_OnAStoreWithoutTimescaleDb_CompletesAsANoOp (RED before the extension guard); probe pins (a store at V150 without the rollup does not read as V152, a TimescaleDB store at V151 reads as V151, and the probe on a migrated store maps to this build); ARunningRollupRefreshHoldingTheMaterializationLockForEightSeconds_IsWaitedOut_AndTheMigrationReaches152 (RED at a 5 s lock timeout).
  • Totals: the V152 set passes with TimescaleDB, and the fresh-database classes pass on a database without it. Darling.Tests and Lite.Tests build with 0 warnings.

Measurements

On the rig, 49,920 query_store_stats rows (10 servers × 24 hours × 208 queries), one hourly rollup of the same shape, refreshed over the day. WAL was measured by pg_wal_lsn_diff around the refresh, twice on independent windows:

run 1 run 2
with the group indexes 32.0 MiB 31.9 MiB
without them 11.2 MiB 11.2 MiB
ratio 2.85× 2.84×

This is a rig number, for one rollup at a small fraction of production scale.

CHANGELOG

SECTION: Changed
ENTRY:

IMPORTANT: Store rung V152 drops unused indexes on six Query Store rollups (no data changes; rows read the same before and after). During a one-time upgrade start that lands on a running rollup refresh, Query Store trend reads can wait up to ~4.5 minutes, and collection starts after the upgrade finishes.

…ame (#4503)

The DO block now resolves each view's materialization OID and matches a
two-key btree by its indexed columns (pg_index/pg_attribute), instead of
building <mat_table>_<column>_bucket_idx as a string and passing it to
DROP INDEX IF EXISTS. PostgreSQL truncates an identifier at 63 bytes with
no hash suffix, so the string-built name for query_store_stats_interval_daily's
runtime_stats_interval_id index was never the name actually on disk, and the
drop would have silently no-op'd.

Also gives every fresh store the same shape an upgraded store reaches: the
six Query Store rollups' Create...Sql now carry
timescaledb.create_group_indexes = false, and EnsureContinuousAggregatesAsync
issues an explicit CREATE INDEX IF NOT EXISTS for each view's kept indexes
(server_id, bucket) plus (server_name, bucket) on the two hourly views right
after the view's own CREATE.
Adds the Viewer probe sentinel and arm for V152 (#4503), the six Query
Store rollups' auto-created group index drop: a new negative sentinel
column that walks each rollup view's own pg_rewrite/pg_depend
dependency to its materialization hypertable rather than querying
timescaledb_information.continuous_aggregates directly, so it needs
no TimescaleDB-extension guard.

Adds CaggGroupIndexDropRungTests (the new top-rung ladder and probe
facts) and rewrites AgGroupIdRungTests to the stays-true shape now
that V151 is no longer the top rung.

Declares V152's DO block (EXECUTE format(...) resolving and dropping
a catalog index by shape) as an exempted dynamic-SQL rung in
MigrationDataMovingRungCensusPins, since it drops rather than builds
or rewrites an index and so needs no data-moving cost declaration of
its own.
…#4503)

A store already at V151, with its six Query Store rollups materialized and
compressed, migrates to V152 and every reader query returns the same rows
before and after the migration. The pin builds each rollup in the pre-V152
shape (TimescaleDB's default group indexes, no kept-index statements),
seeds two servers across two days, refreshes and compresses every
materialization the product's own way, captures every named reader's rows,
runs the V152 rung, and re-captures. It then checks the index shape
directly: the kept two-key indexes are present and the dropped ones are
gone, matched by key columns rather than by name.

New file only: Darling/Darling.Tests/CaggGroupIndexDropUpgradeLiveTests.cs.
…re keeps (#4503)

Each Query Store rollup's kept group index (server_id/bucket, and
server_name/bucket on the two hourly views) is now created unnamed,
directly on the continuous aggregate's view, in bucket DESC order --
the same shape TimescaleDB's own create_group_indexes builds and the
same shape an upgraded store's surviving auto-created index carries.
PostgreSQL assigns the index its own default name from the
materialization's real (possibly truncated) table name, so a fresh
store and an upgraded store land byte-identical pg_indexes rows: same
key order, same DESC sort, same name.

Previously the fresh-store statements built an explicit name and used
ASC order, which made the fresh side's index a different physical
object from the survivor an upgraded store keeps after V152's DO
block runs.

Extends PgSchemaGeneratorTests with a static text pin comparing the
fresh-side SQL against V152's catalog-shape drop logic for all six
views, and adds CaggGroupIndexDropLiveTests: a live pin that a fresh
migrate plus the aggregate sweep keeps only the surviving group
indexes on all six materializations (RED at runtime against
origin/dev, where the dropped indexes still exist), and a live pin
that running the aggregate sweep twice is idempotent on those kept
indexes.
)

Three tests pinned the pre-V152 shape of the six Query Store rollups'
continuous-aggregate SQL and failed once the rung gave them
timescaledb.create_group_indexes = false:

- TimescaleContinuousAggregateTests.QueryStoreStatsHourly_GroupedByComposerDims_CarriesWeightedSums
  and ...QueryStoreStatsIntervalDaily_RededupsTheIntervalAcrossTheDay_FromL1
  asserted the bare WITH (timescaledb.continuous), with no option. Updated
  to assert the option is present, with a comment citing the production
  catalog read that measured it.
- IntervalDedupMaterializationIndexesTests.EveryOtherAggregate_KeepsTheDefaultGroupIndexes_UntilMeasured
  is the scope pin: every aggregate NOT on the measured list keeps
  TimescaleDB's default. Widened to exclude the six newly-measured
  rollups from the exempt set, and to assert the other direction too —
  those six DO carry the option, not silently dropped from the pin.

All three fail on origin/dev (which lacks the option) and pass at this
branch's head; confirmed both ways before rewriting.
…r probe false positive, add a lock_timeout, and tighten the drop match

V152 unconditionally queried timescaledb_information.continuous_aggregates,
which raises 42P01 on a store that never created the TimescaleDB extension.
The rung's DO block now checks to_regclass('timescaledb_information.continuous_aggregates')
first and returns as a no-op when it is absent, the same guard pattern the
baseline-fallback-view DO blocks already use. The per-view catalog read now
goes through EXECUTE ... INTO ... USING so the timescaledb_information
reference is parsed only when the guard has already confirmed it exists.

The Viewer's schema-version probe for V152 read as a false positive: its
bare NOT EXISTS over the dependency walk is also true on any store missing
the query_store_stats_hourly rollup (a plain-PostgreSQL store, or a
TimescaleDB store whose rollup isn't materialized yet), which would report
that store as fully current at V152 regardless of its real version. The
sentinel now ANDs in a positive existence check on the view first, the same
shape V149's negative probe already uses.

DROP INDEX takes an ACCESS EXCLUSIVE lock on the materialization hypertable
and every one of its chunks, across six hypertables in one migration
transaction, while a background refresh policy can hold the same lock for
minutes. V152 now runs under SET LOCAL lock_timeout = '5s' so a live refresh
makes the rung fail (retryable) rather than queue for the whole migration
command timeout.

The drop match also now excludes UNIQUE indexes and INCLUDE columns
(NOT indisunique AND indnkeyatts = 2), so an operator-added unique
(query_hash, bucket) index is never mistaken for the auto-created group
index. A handful of doc comments were corrected to match what the code
actually does: the fresh-migrate kept-index build's idempotency guard is a
catalog read, not IF NOT EXISTS; the kept index is created against the
materialization's OID directly, not against the continuous aggregate's
view; corrected_hourly's kept server_name index is justified by scoped
composer reads (it measured 0 lifetime scans, not the 131 the plain hourly
view's index showed); and the dynamic-SQL census exemption for V152 now
names both EXECUTE forms the rung uses.

Adds a new live pin, V152_OnAStoreWithoutTimescaleDb_CompletesAsANoOp: a
plain migrate that never enables TimescaleDB reaches the top with no
exception and the schema version recorded as StorageVersion.SchemaVersion.

Refs #4503
The V152 sentinel ANDed a to_regclass(...) IS NOT NULL existence check
on query_store_stats_hourly with its NOT EXISTS group-index absence
check. A plain-PostgreSQL store has no rollup at all, so to_regclass
resolves to NULL there at every version, the dependency walk finds no
rows, and the bare NOT EXISTS read that absence as "index dropped" -
misreporting a plain-PostgreSQL store at V151 as V152.

The sentinel now ANDs the NOT EXISTS check with V151's own sentinel
(the ag_replica_states.group_id column), copied verbatim, instead of
the existence check on the view. A store below V151 fails that
sentinel and falls through to the next arm; a plain-PostgreSQL store
at V151 now correctly reads as V152 too, which is harmless because
V152 is a no-op without TimescaleDB, so the two schemas are identical
there.

Adds two live pins: a V150 store with no rollup does not map to 152,
and a TimescaleDB store at V151 (group index still present) maps to
exactly 151.
…imeout

V152's SET LOCAL lock_timeout was a flat 5s, shorter than a running hourly
Query Store rollup refresh can hold the AccessExclusiveLock its DROP INDEX
needs. Set it to MigrationCommandTimeoutSeconds minus a 20s margin (280s)
instead, so the rung waits out one whole refresh cycle while the server's
own retryable 55P03 still fires before the client's command timeout would
cancel the statement out from under it.

Added a live pin: a second connection holds an ACCESS SHARE lock on a
V151-shape rollup's materialization for about 8s while the migration runs
on another connection; it must finish once the lock releases rather than
fail. That pin fails with 55P03 against the previous flat 5s value.

Refs #4503
@erikdarlingdata
erikdarlingdata marked this pull request as ready for review September 28, 2026 00:46
@erikdarlingdata
erikdarlingdata merged commit 16efefb into dev Sep 28, 2026
17 of 18 checks passed
@erikdarlingdata
erikdarlingdata deleted the feat/v152-drop-unread-cagg-group-indexes branch September 28, 2026 00:46
erikdarlingdata added a commit that referenced this pull request Sep 28, 2026
…00Z UTC (#4621)

Moves the CHANGELOG entries carried in merged pull-request descriptions into [Unreleased]. The cut is PRs merged after 2026-09-26T17:37:33Z and at or before 2026-09-28T17:40:00Z; the next splice starts after it.

- 76 PRs are spliced: Fixed 36, Changed 28 and Added 12, each counted once under its first section. That is 79 bullets: #4481's Fixed entry holds three, and #4548 adds a second group under Added.
- 21 PRs have no user-visible entry (None, test-only, CI-only, or deferred to a parent).
- The [#4509] link definition, which #4538 also carries, is defined once.
- The IMPORTANT upgrade notes from #4489, #4495, #4501, #4506 and #4541 are held for the release cut and are not in this change.
- Only CHANGELOG.md changes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant