Skip to content

procedure_stats is 9.2x slower on use1 than use2 for identical row counts, and is 69% of use1's collection body #2847

Description

@erikdarlingdata

procedure_stats takes 9.2x longer on use1 than use2 for identical row counts

Split out of #2841, where it turned out to be the dominant term. Measured from collect.collection_log on
both stores, window >= 2026-09-03 14:00 UTC — after the t3.xlarge -> m7i.2xlarge resize, so both boxes
are unthrottled 8 vCPU / 31 GB for the whole window. 42 servers each.

use1 use2 ratio
duration_ms p50 4,900 535 9.2x
duration_ms p90 7,161 1,019 7.0x
sql_duration_ms p50 4,635 473 9.8x
duckdb_duration_ms p50 224 58 3.9x
median rows_collected 150 150 1.0x
n 2,481 4,200

It is not workload volume, and it is not a few bad targets

Same row count. Median rows_collected is 150 on both boxes, and the top-12 slowest use1 servers all
return exactly 150. So it is ~33 ms per row on use1 against ~3.5 ms on use2 for the same output.

Uniform across the fleet. Per-server p50 on use1: median 4,952 ms, p90 6,154 ms, worst
7,532 ms, across all 42 servers. No outliers carrying it — every server is slow.

Target-side. sql_duration_ms is 95% of duration_ms on use1 (4,635 of 4,900), so it is the query
against the monitored server, not the store write.

Why it matters beyond this collector

At p50 4,900 ms it is 69% of use1's entire 1-minute collection body (7,061 ms summed across the 19
collectors on that tier). Excluding it, use1's body is 2,161 ms against use2's 1,665 ms — 1.3x apart rather
than 3.2x. It is the single reason the two boxes' bodies differ, and therefore the largest lever on #2841's
cadence floor.

Candidate causes — none verified

The obvious environmental difference is that Query Store is live on use1's targets and dead on use2's
since the #2296 split (2026-08-17). Whether procedure_stats touches anything Query-Store-adjacent, or
whether use1's targets simply carry larger plan caches as a consequence, is unestablished. Other candidates:
per-row plan or text lookups that scale with cache size, or a genuine difference in the monitored workloads
that rows_collected does not capture.

The decisive diagnostic is the collector's actual execution plan captured on a use1 target and a use2
target with the same shipped query text — not a retyped copy. That splits "the query is doing more work
because the target has more to look at" from "the query has a plan problem on one fleet".

Constraint on the fix

Rescheduling it off the 1-minute tier is not freely available: PR #2843 (0ebbc7db) added a pin
asserting no detached collector sits on the 1-minute tier, because a detached run skips when its previous
run is still in flight — so detaching a 1-minute collector converts starvation into guaranteed misses.
Changing its default frequency instead is a CollectorScheduleDefaults change, which is duplicated into
Lite's ScheduleManager.s_presets and needs the parity treatment (and Lite.Tests does not run in PR CI).

Note on configuration

config.config_collector_schedules is empty on both stores — zero rows. Neither box has per-server
schedule overrides, so both run the shared CollectorScheduleDefaults cadences directly. Any comparison
between the two boxes is therefore comparing identical schedules, and any fix to the default frequency
lands on both.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions