Summary
PgBaselineProvider sets no CommandTimeout, so its baseline query inherits Npgsql's undocumented 30 s default — a value nobody chose. This is the same defect #2810 found and fixed in PgFactCollector, in the other half of the same 120 s analysis pass.
It is failing in production today.
Measured evidence
Every observed failure sits on a hard 30 s wall. From the dogfood box's app log, 19 occurrences across two days:
| when (UTC) |
metric |
elapsed |
| 2026-09-03 00:27:54 |
io_latency |
31.4 s |
| 2026-09-03 02:41:28 |
io_latency |
30.4 s |
| 2026-09-03 02:41:58 |
io_latency |
30.2 s |
| 2026-09-03 02:42:22 |
io_latency |
30.3 s |
| 2026-09-03 02:42:43 |
io_latency |
30.3 s |
| 2026-09-03 02:43:21 |
io_latency |
30.3 s |
| 2026-09-03 04:14:56 |
io_latency |
30.7 s |
| 2026-09-03 04:16:14 |
io_latency |
31.1 s |
| … |
|
|
| 2026-09-02 23:25:52 |
io_latency |
30.7 s |
No sample falls outside 30.1–31.4 s. That is not a variable-duration failure and it is not the 15 s connection timeout — it is a fixed ceiling being hit.
The query itself is fast. Running the shipped query string (dumped from the built assembly, not retyped) against the live store, three busiest servers by file_io_stats row count in the 30-day window:
| store |
first execution |
| use2 srv A |
1,624 ms |
| use2 srv B |
1,476 ms |
| use2 srv C |
1,362 ms |
| use1 srv A |
230 ms |
| use1 srv B |
336 ms |
| use1 srv C |
198 ms |
So the read normally completes in ~1.6 s worst case and only crosses 30 s during bursts of store-wide slowness — note the clustering above, five failures inside two minutes and five more inside four. Exactly #2810's shape: it fails on the bad combination only, which is why it reads as intermittent.
The failure data is right-censored. Every failing run was killed at 30 s, so nothing in the record says whether it needed 35 s or 300 s. Any bound has to be chosen against that limitation rather than pretending to a number the data does not contain.
Why the two previous changes could not have fixed it
Neither touched the ceiling. A query at 4.2 s steady-state still dies at 30 s when the store stalls, which is what the log shows.
The cost when it fires
Per the collector's own message: that metric has no baseline for the pass and its anomaly detection is silent. Worse, a null result is cached — _cache[cacheKey] is assigned unconditionally in GetOrComputeBaselinesAsync — so a single stall silences that (server, metric) pair for the full 1-hour CacheTtl, not just the pass that hit it.
Scope
Sweeping by shape rather than by the one reported site, as #2810 did with its 31. The analysis pass has 32 remaining NpgsqlCommand sites with no deadline:
| file |
untimed sites |
PgAnomalyDetector.cs |
12 |
PgFindingStore.cs |
8 |
PgDrillDownCollector.Blocking.cs |
4 |
PgDrillDownCollector.Storage.cs |
3 |
PgDrillDownCollector.Plans.cs |
2 |
PgBaselineProvider.cs |
1 |
PgDrillDownCollector.Config.cs |
1 |
DarlingAnalysisService.cs |
1 |
All share the 120 s DarlingWorker.s_analysisTimeout budget and all had the same missing value. Pinning one while leaving thirty-one identical ones is the enumerated-list mistake this repo has paid for repeatedly.
Note PgDrillDownCollector.Queries.cs already carries a deliberate DrillDownCommandTimeoutSeconds = 30 on 8 of that collector's sites; its other 10 will get that same existing constant rather than a new value, so an author's existing choice is not silently overridden.
Remaining after this change
Repo-wide (Darling, non-test): 275 NpgsqlCommand sites, 110 with a deadline, 165 without. This closes the 32 in the analysis pass, leaving 133 — 111 in .Service, 20 in .Storage, 2 in .Viewer. Those sit outside the analysis pass and have different budgets, so they warrant their own issue rather than being swept in here.
Summary
PgBaselineProvidersets noCommandTimeout, so its baseline query inherits Npgsql's undocumented 30 s default — a value nobody chose. This is the same defect #2810 found and fixed inPgFactCollector, in the other half of the same 120 s analysis pass.It is failing in production today.
Measured evidence
Every observed failure sits on a hard 30 s wall. From the dogfood box's app log, 19 occurrences across two days:
No sample falls outside 30.1–31.4 s. That is not a variable-duration failure and it is not the 15 s connection timeout — it is a fixed ceiling being hit.
The query itself is fast. Running the shipped query string (dumped from the built assembly, not retyped) against the live store, three busiest servers by
file_io_statsrow count in the 30-day window:So the read normally completes in ~1.6 s worst case and only crosses 30 s during bursts of store-wide slowness — note the clustering above, five failures inside two minutes and five more inside four. Exactly #2810's shape: it fails on the bad combination only, which is why it reads as intermittent.
The failure data is right-censored. Every failing run was killed at 30 s, so nothing in the record says whether it needed 35 s or 300 s. Any bound has to be chosen against that limitation rather than pretending to a number the data does not contain.
Why the two previous changes could not have fixed it
UNION ALL+ equality.Neither touched the ceiling. A query at 4.2 s steady-state still dies at 30 s when the store stalls, which is what the log shows.
The cost when it fires
Per the collector's own message: that metric has no baseline for the pass and its anomaly detection is silent. Worse, a null result is cached —
_cache[cacheKey]is assigned unconditionally inGetOrComputeBaselinesAsync— so a single stall silences that (server, metric) pair for the full 1-hourCacheTtl, not just the pass that hit it.Scope
Sweeping by shape rather than by the one reported site, as #2810 did with its 31. The analysis pass has 32 remaining
NpgsqlCommandsites with no deadline:PgAnomalyDetector.csPgFindingStore.csPgDrillDownCollector.Blocking.csPgDrillDownCollector.Storage.csPgDrillDownCollector.Plans.csPgBaselineProvider.csPgDrillDownCollector.Config.csDarlingAnalysisService.csAll share the 120 s
DarlingWorker.s_analysisTimeoutbudget and all had the same missing value. Pinning one while leaving thirty-one identical ones is the enumerated-list mistake this repo has paid for repeatedly.Note
PgDrillDownCollector.Queries.csalready carries a deliberateDrillDownCommandTimeoutSeconds = 30on 8 of that collector's sites; its other 10 will get that same existing constant rather than a new value, so an author's existing choice is not silently overridden.Remaining after this change
Repo-wide (Darling, non-test): 275
NpgsqlCommandsites, 110 with a deadline, 165 without. This closes the 32 in the analysis pass, leaving 133 — 111 in.Service, 20 in.Storage, 2 in.Viewer. Those sit outside the analysis pass and have different budgets, so they warrant their own issue rather than being swept in here.