procedure_stats takes 9.2x longer on use1 than use2 for identical row counts
Split out of #2841, where it turned out to be the dominant term. Measured from collect.collection_log on
both stores, window >= 2026-09-03 14:00 UTC — after the t3.xlarge -> m7i.2xlarge resize, so both boxes
are unthrottled 8 vCPU / 31 GB for the whole window. 42 servers each.
|
use1 |
use2 |
ratio |
duration_ms p50 |
4,900 |
535 |
9.2x |
duration_ms p90 |
7,161 |
1,019 |
7.0x |
sql_duration_ms p50 |
4,635 |
473 |
9.8x |
duckdb_duration_ms p50 |
224 |
58 |
3.9x |
median rows_collected |
150 |
150 |
1.0x |
| n |
2,481 |
4,200 |
|
It is not workload volume, and it is not a few bad targets
Same row count. Median rows_collected is 150 on both boxes, and the top-12 slowest use1 servers all
return exactly 150. So it is ~33 ms per row on use1 against ~3.5 ms on use2 for the same output.
Uniform across the fleet. Per-server p50 on use1: median 4,952 ms, p90 6,154 ms, worst
7,532 ms, across all 42 servers. No outliers carrying it — every server is slow.
Target-side. sql_duration_ms is 95% of duration_ms on use1 (4,635 of 4,900), so it is the query
against the monitored server, not the store write.
Why it matters beyond this collector
At p50 4,900 ms it is 69% of use1's entire 1-minute collection body (7,061 ms summed across the 19
collectors on that tier). Excluding it, use1's body is 2,161 ms against use2's 1,665 ms — 1.3x apart rather
than 3.2x. It is the single reason the two boxes' bodies differ, and therefore the largest lever on #2841's
cadence floor.
Candidate causes — none verified
The obvious environmental difference is that Query Store is live on use1's targets and dead on use2's
since the #2296 split (2026-08-17). Whether procedure_stats touches anything Query-Store-adjacent, or
whether use1's targets simply carry larger plan caches as a consequence, is unestablished. Other candidates:
per-row plan or text lookups that scale with cache size, or a genuine difference in the monitored workloads
that rows_collected does not capture.
The decisive diagnostic is the collector's actual execution plan captured on a use1 target and a use2
target with the same shipped query text — not a retyped copy. That splits "the query is doing more work
because the target has more to look at" from "the query has a plan problem on one fleet".
Constraint on the fix
Rescheduling it off the 1-minute tier is not freely available: PR #2843 (0ebbc7db) added a pin
asserting no detached collector sits on the 1-minute tier, because a detached run skips when its previous
run is still in flight — so detaching a 1-minute collector converts starvation into guaranteed misses.
Changing its default frequency instead is a CollectorScheduleDefaults change, which is duplicated into
Lite's ScheduleManager.s_presets and needs the parity treatment (and Lite.Tests does not run in PR CI).
Note on configuration
config.config_collector_schedules is empty on both stores — zero rows. Neither box has per-server
schedule overrides, so both run the shared CollectorScheduleDefaults cadences directly. Any comparison
between the two boxes is therefore comparing identical schedules, and any fix to the default frequency
lands on both.
procedure_statstakes 9.2x longer on use1 than use2 for identical row countsSplit out of #2841, where it turned out to be the dominant term. Measured from
collect.collection_logonboth stores, window >= 2026-09-03 14:00 UTC — after the t3.xlarge -> m7i.2xlarge resize, so both boxes
are unthrottled 8 vCPU / 31 GB for the whole window. 42 servers each.
duration_msp50duration_msp90sql_duration_msp50duckdb_duration_msp50rows_collectedIt is not workload volume, and it is not a few bad targets
Same row count. Median
rows_collectedis 150 on both boxes, and the top-12 slowest use1 servers allreturn exactly 150. So it is ~33 ms per row on use1 against ~3.5 ms on use2 for the same output.
Uniform across the fleet. Per-server p50 on use1: median 4,952 ms, p90 6,154 ms, worst
7,532 ms, across all 42 servers. No outliers carrying it — every server is slow.
Target-side.
sql_duration_msis 95% ofduration_mson use1 (4,635 of 4,900), so it is the queryagainst the monitored server, not the store write.
Why it matters beyond this collector
At p50 4,900 ms it is 69% of use1's entire 1-minute collection body (7,061 ms summed across the 19
collectors on that tier). Excluding it, use1's body is 2,161 ms against use2's 1,665 ms — 1.3x apart rather
than 3.2x. It is the single reason the two boxes' bodies differ, and therefore the largest lever on #2841's
cadence floor.
Candidate causes — none verified
The obvious environmental difference is that Query Store is live on use1's targets and dead on use2's
since the #2296 split (2026-08-17). Whether
procedure_statstouches anything Query-Store-adjacent, orwhether use1's targets simply carry larger plan caches as a consequence, is unestablished. Other candidates:
per-row plan or text lookups that scale with cache size, or a genuine difference in the monitored workloads
that
rows_collecteddoes not capture.The decisive diagnostic is the collector's actual execution plan captured on a use1 target and a use2
target with the same shipped query text — not a retyped copy. That splits "the query is doing more work
because the target has more to look at" from "the query has a plan problem on one fleet".
Constraint on the fix
Rescheduling it off the 1-minute tier is not freely available: PR #2843 (
0ebbc7db) added a pinasserting no detached collector sits on the 1-minute tier, because a detached run skips when its previous
run is still in flight — so detaching a 1-minute collector converts starvation into guaranteed misses.
Changing its default frequency instead is a
CollectorScheduleDefaultschange, which is duplicated intoLite's
ScheduleManager.s_presetsand needs the parity treatment (andLite.Testsdoes not run in PR CI).Note on configuration
config.config_collector_schedulesis empty on both stores — zero rows. Neither box has per-serverschedule overrides, so both run the shared
CollectorScheduleDefaultscadences directly. Any comparisonbetween the two boxes is therefore comparing identical schedules, and any fix to the default frequency
lands on both.