#2617 closed on rig evidence (160 indexes, 56ms) with its own real-cluster acceptance criteria left for "tomorrow's run": SUCCESS, 1,517 rows, ~200 measured, largest first. Nobody checked tomorrow's run. The store has now answered it eleven times.
Live state (pgmon store, read 2026-09-05)
Every pg_index_bloat run that has ever executed on the #2617 target (target-a) is status=ERROR, rows=0, message Exception while reading from stream, daily from the first-ever run at 2026-08-25T18:14:21Z (the run #2617 quotes) through 2026-09-04T19:00:27Z — eleven for eleven, including all nine runs after #2618's fix was hash-verified onto this box (#2617's close comment records the deploy at 2026-08-26T00:24Z; every subsequent daily run still fails identically). last_success: null.
The other 49 pgmon targets are all NO_PERMISSIONS (pgstatindex absent — pgstattuple never created), last_success: null. So the collector has produced zero rows fleet-wide, ever, and #2617's own closing worry stands: the #2561 exact-bloat claim remains unverified against any production cluster.
Cost on the failing target: the run is not free. Around the 2026-08-30 failure the entire sweep froze for 316 seconds (18:29:35 → 18:34:51), every 1-minute collector gapping, resuming the second the ERROR landed — and the daily run time drifts ~5 minutes later each day, which is a ~300-second slot being consumed daily. That is the collector blocking the sweep body for its full timeout before dying.
Mechanism, two layers — both verified in source on dev
1. The work budget bounds the wrong dimension for this target. PgIndexBloatCollector.cs:176-181 chose the 200-index count deliberately — "200 rather than a byte budget … A count is legible" — with MeasureCeilingBytes = 20GB (line 63) bounding any single index and largest-first ordering choosing which 200. On target-a that legibility trade fails: reading just the top 80 indexes from the store, 3 exceed the ceiling (correctly excluded, 71GB) and 77 sub-ceiling indexes sum 286GB, topping out at 19.68GB each — an independent read by another lane got the same shape (57 indexes / 265GB under a slightly different cut). The 200 largest sub-ceiling indexes are therefore hundreds of GB of pgstatindex page reads in one LEFT JOIN LATERAL statement. The 300-second CommandTimeoutSecondsOverride (line 190) cannot cover that; the count budget only bounds work when count correlates with bytes, and this target is the counterexample. #2617's own diagnosis — "a per-index ceiling and a total-work ceiling are different things and I only built the first" — happened a second time, one dimension up.
2. The timeout the fix added can never be classified. #2618's remedy 3 wanted "a classified timeout rather than a dropped connection". But the PostgreSQL classification arm in DarlingWorker.cs is:
catch (PostgresException ex) when (
PostgresFaultOutcome(ex, collectorName, runtime.ConnectedDatabase) is { Status: not "ERROR" } outcome)
A client-side command timeout surfaces as NpgsqlException wrapping TimeoutException — not PostgresException, no SQLSTATE — so it can never enter PostgresFaultOutcome and falls to the general catch, which writes "ERROR", 0, 0, 0 with null phase blocks and the raw seven-word transport message. Two side effects worth naming: the hardcoded zeros make a 300-second death look like an instant one (it misled two independent investigations toward "connection dies at open"), and the message carries no collector/database/duration context at all — compare the authored text every classified skip on this fleet gets.
Fix shape, in this order (order matters)
- Widen the PG fault classification to cover non-
PostgresException faults — an authored CommandTimeout outcome ("client-side deadline of Ns expired; the statement was cancelled mid-read") and a ConnectionFatal one — without disturbing the existing SQLSTATE path that correctly degrades the other 49 targets to PERMISSIONS. Doing 2 first would leave the next timeout as unreadable as this one.
- Bound bytes, not just count: cap the cycle at a cumulative sub-ceiling byte budget (keep largest-first; keep
skipped_reason = not measured this cycle (work budget) for the remainder). The :176 comment's legibility argument can survive as "N indexes or B bytes, whichever first".
- Record real elapsed time on the ERROR arm instead of literal zeros — the current shape destroys exactly the evidence that distinguishes a timeout from a connection failure.
The observation that decides
Next scheduled attempt on this target is ~2026-09-05T19:05Z (daily cadence drifting from 18:14 → 19:00 across the eleven runs — note it is NOT in the ~10:33Z block the other dailies use), and it will be the first attempt on the build deployed 2026-09-05 04:13Z. Expected under this diagnosis: identical failure, since nothing in that build touches the budget or the classification. A SUCCESS there falsifies layer 1 and leaves only layer 2 to fix.
Related: #2617 (closed on the rig), #2618 (the fix this measures), #2561 (the claim still unverified in production), #2994 (the same rollup-honesty family on this fleet), #2623.
#2617 closed on rig evidence (160 indexes, 56ms) with its own real-cluster acceptance criteria left for "tomorrow's run": SUCCESS, 1,517 rows, ~200 measured, largest first. Nobody checked tomorrow's run. The store has now answered it eleven times.
Live state (pgmon store, read 2026-09-05)
Every
pg_index_bloatrun that has ever executed on the #2617 target (target-a) isstatus=ERROR,rows=0, messageException while reading from stream, daily from the first-ever run at 2026-08-25T18:14:21Z (the run #2617 quotes) through 2026-09-04T19:00:27Z — eleven for eleven, including all nine runs after #2618's fix was hash-verified onto this box (#2617's close comment records the deploy at 2026-08-26T00:24Z; every subsequent daily run still fails identically).last_success: null.The other 49 pgmon targets are all
NO_PERMISSIONS(pgstatindexabsent — pgstattuple never created),last_success: null. So the collector has produced zero rows fleet-wide, ever, and #2617's own closing worry stands: the #2561 exact-bloat claim remains unverified against any production cluster.Cost on the failing target: the run is not free. Around the 2026-08-30 failure the entire sweep froze for 316 seconds (18:29:35 → 18:34:51), every 1-minute collector gapping, resuming the second the ERROR landed — and the daily run time drifts ~5 minutes later each day, which is a ~300-second slot being consumed daily. That is the collector blocking the sweep body for its full timeout before dying.
Mechanism, two layers — both verified in source on
dev1. The work budget bounds the wrong dimension for this target.
PgIndexBloatCollector.cs:176-181chose the 200-index count deliberately — "200 rather than a byte budget … A count is legible" — withMeasureCeilingBytes = 20GB(line 63) bounding any single index and largest-first ordering choosing which 200. On target-a that legibility trade fails: reading just the top 80 indexes from the store, 3 exceed the ceiling (correctly excluded, 71GB) and 77 sub-ceiling indexes sum 286GB, topping out at 19.68GB each — an independent read by another lane got the same shape (57 indexes / 265GB under a slightly different cut). The 200 largest sub-ceiling indexes are therefore hundreds of GB ofpgstatindexpage reads in oneLEFT JOIN LATERALstatement. The 300-secondCommandTimeoutSecondsOverride(line 190) cannot cover that; the count budget only bounds work when count correlates with bytes, and this target is the counterexample. #2617's own diagnosis — "a per-index ceiling and a total-work ceiling are different things and I only built the first" — happened a second time, one dimension up.2. The timeout the fix added can never be classified. #2618's remedy 3 wanted "a classified timeout rather than a dropped connection". But the PostgreSQL classification arm in
DarlingWorker.csis:A client-side command timeout surfaces as
NpgsqlExceptionwrappingTimeoutException— notPostgresException, no SQLSTATE — so it can never enterPostgresFaultOutcomeand falls to the general catch, which writes"ERROR", 0, 0, 0with null phase blocks and the raw seven-word transport message. Two side effects worth naming: the hardcoded zeros make a 300-second death look like an instant one (it misled two independent investigations toward "connection dies at open"), and the message carries no collector/database/duration context at all — compare the authored text every classified skip on this fleet gets.Fix shape, in this order (order matters)
PostgresExceptionfaults — an authoredCommandTimeoutoutcome ("client-side deadline of Ns expired; the statement was cancelled mid-read") and aConnectionFatalone — without disturbing the existing SQLSTATE path that correctly degrades the other 49 targets to PERMISSIONS. Doing 2 first would leave the next timeout as unreadable as this one.skipped_reason = not measured this cycle (work budget)for the remainder). The :176 comment's legibility argument can survive as "N indexes or B bytes, whichever first".The observation that decides
Next scheduled attempt on this target is ~2026-09-05T19:05Z (daily cadence drifting from 18:14 → 19:00 across the eleven runs — note it is NOT in the ~10:33Z block the other dailies use), and it will be the first attempt on the build deployed 2026-09-05 04:13Z. Expected under this diagnosis: identical failure, since nothing in that build touches the budget or the classification. A SUCCESS there falsifies layer 1 and leaves only layer 2 to fix.
Related: #2617 (closed on the rig), #2618 (the fix this measures), #2561 (the claim still unverified in production), #2994 (the same rollup-honesty family on this fleet), #2623.