Skip to content

pg_index_bloat has failed every run since #2618 shipped: the 200-index budget admits 286GB on target-a, and the timeout it dies on cannot be classified #2997

Description

@erikdarlingdata

#2617 closed on rig evidence (160 indexes, 56ms) with its own real-cluster acceptance criteria left for "tomorrow's run": SUCCESS, 1,517 rows, ~200 measured, largest first. Nobody checked tomorrow's run. The store has now answered it eleven times.

Live state (pgmon store, read 2026-09-05)

Every pg_index_bloat run that has ever executed on the #2617 target (target-a) is status=ERROR, rows=0, message Exception while reading from stream, daily from the first-ever run at 2026-08-25T18:14:21Z (the run #2617 quotes) through 2026-09-04T19:00:27Z — eleven for eleven, including all nine runs after #2618's fix was hash-verified onto this box (#2617's close comment records the deploy at 2026-08-26T00:24Z; every subsequent daily run still fails identically). last_success: null.

The other 49 pgmon targets are all NO_PERMISSIONS (pgstatindex absent — pgstattuple never created), last_success: null. So the collector has produced zero rows fleet-wide, ever, and #2617's own closing worry stands: the #2561 exact-bloat claim remains unverified against any production cluster.

Cost on the failing target: the run is not free. Around the 2026-08-30 failure the entire sweep froze for 316 seconds (18:29:35 → 18:34:51), every 1-minute collector gapping, resuming the second the ERROR landed — and the daily run time drifts ~5 minutes later each day, which is a ~300-second slot being consumed daily. That is the collector blocking the sweep body for its full timeout before dying.

Mechanism, two layers — both verified in source on dev

1. The work budget bounds the wrong dimension for this target. PgIndexBloatCollector.cs:176-181 chose the 200-index count deliberately — "200 rather than a byte budget … A count is legible" — with MeasureCeilingBytes = 20GB (line 63) bounding any single index and largest-first ordering choosing which 200. On target-a that legibility trade fails: reading just the top 80 indexes from the store, 3 exceed the ceiling (correctly excluded, 71GB) and 77 sub-ceiling indexes sum 286GB, topping out at 19.68GB each — an independent read by another lane got the same shape (57 indexes / 265GB under a slightly different cut). The 200 largest sub-ceiling indexes are therefore hundreds of GB of pgstatindex page reads in one LEFT JOIN LATERAL statement. The 300-second CommandTimeoutSecondsOverride (line 190) cannot cover that; the count budget only bounds work when count correlates with bytes, and this target is the counterexample. #2617's own diagnosis — "a per-index ceiling and a total-work ceiling are different things and I only built the first" — happened a second time, one dimension up.

2. The timeout the fix added can never be classified. #2618's remedy 3 wanted "a classified timeout rather than a dropped connection". But the PostgreSQL classification arm in DarlingWorker.cs is:

catch (PostgresException ex) when (
    PostgresFaultOutcome(ex, collectorName, runtime.ConnectedDatabase) is { Status: not "ERROR" } outcome)

A client-side command timeout surfaces as NpgsqlException wrapping TimeoutException — not PostgresException, no SQLSTATE — so it can never enter PostgresFaultOutcome and falls to the general catch, which writes "ERROR", 0, 0, 0 with null phase blocks and the raw seven-word transport message. Two side effects worth naming: the hardcoded zeros make a 300-second death look like an instant one (it misled two independent investigations toward "connection dies at open"), and the message carries no collector/database/duration context at all — compare the authored text every classified skip on this fleet gets.

Fix shape, in this order (order matters)

  1. Widen the PG fault classification to cover non-PostgresException faults — an authored CommandTimeout outcome ("client-side deadline of Ns expired; the statement was cancelled mid-read") and a ConnectionFatal one — without disturbing the existing SQLSTATE path that correctly degrades the other 49 targets to PERMISSIONS. Doing 2 first would leave the next timeout as unreadable as this one.
  2. Bound bytes, not just count: cap the cycle at a cumulative sub-ceiling byte budget (keep largest-first; keep skipped_reason = not measured this cycle (work budget) for the remainder). The :176 comment's legibility argument can survive as "N indexes or B bytes, whichever first".
  3. Record real elapsed time on the ERROR arm instead of literal zeros — the current shape destroys exactly the evidence that distinguishes a timeout from a connection failure.

The observation that decides

Next scheduled attempt on this target is ~2026-09-05T19:05Z (daily cadence drifting from 18:14 → 19:00 across the eleven runs — note it is NOT in the ~10:33Z block the other dailies use), and it will be the first attempt on the build deployed 2026-09-05 04:13Z. Expected under this diagnosis: identical failure, since nothing in that build touches the budget or the classification. A SUCCESS there falsifies layer 1 and leaves only layer 2 to fix.

Related: #2617 (closed on the rig), #2618 (the fix this measures), #2561 (the claim still unverified in production), #2994 (the same rollup-honesty family on this fleet), #2623.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions