Skip to content

Collection cadence has no fixed component: cadence = N x gate_held_time / max_concurrent_sweeps + tick/2 #2849

Description

@erikdarlingdata

Summary

The "~55 s fixed component" in #2841 does not exist. Cadence is fully explained by a purely multiplicative model derived from the launch loop, with no intercept beyond half a sweep tick. Measured on both boxes after today's resize, the model fits use1 to 0.4%.

This reframes #2841 and promotes #2847.

The mechanism (from DarlingWorker launch loop)

The fire-and-track loop (#1553) launches every server every 15 s tick unless its previous body is still in flight. The concurrency gate is acquired inside the body (await gate.WaitAsync(stoppingToken) is the first statement of ProcessServerSweepAsync), so all 42 bodies launch and queue on the semaphore. It is a queue, not a round-robin.

Steady state therefore gives:

cadence ≈ (N × gate_held_time) / C + tick/2

where N = enabled servers, C = max_concurrent_sweeps, tick = s_sweepInterval (15 s). The tick/2 is the average wait for the next launch tick after a body completes.

Measurements (window >= 2026-09-03 14:00Z, both boxes m7i.2xlarge, post-resize)

use1 use2
config_service.max_concurrent_sweeps 4 4
enabled servers 42 42 (61 registered, 19 disabled)
body span p50 / mean (collection_log first→last row per sweep) 8.39 / 9.77 s 2.59 / 3.80 s
collectors per sweep p50 19 19
sweep gap p50 (n=2,750 / 5,081) 109.7 s 69.8 s
sweep gap p90 164.8 s 77.4 s
episode duration p50 (app log, incl. queue wait, n=2,685) 105 s n/a (none surfaced)

Model fit

Solving for gate-held time from observed cadence:

implied R = (cadence − 7.5) × C / N measured body span mean residual
use1 9.73 s 9.77 s 0.4%
use2 5.93 s 3.80 s 2.13 s

use1 fits almost exactly. use2's 2.13 s residual is gate-held time that produces no collection_log row and so falls outside a span measured between first and last row — consistent with the honest-empties convention (a gated-off collector gets no row at all, and use2's Query Store is dead by design since #2296).

Why #2841 saw a 55 s intercept

It fitted a straight line through two points whose x-values were synthetic sums of the 1-minute tier (7,061 ms and 2,200 ms), not measured gate-held time. Measured gate-held time is 9.73 s and 5.93 s — ratio 1.64, versus the synthetic ratio of 3.21. Fitting a line through mis-scaled x-values manufactures an intercept. There is no fixed term.

max_concurrent_sweeps reconciliation

The knob is at its documented default of 4 on both stores — it is not set high, and the limit is being enforced. The earlier "10 and 23 distinct servers per 15 s bucket" observation is not evidence of concurrency above 4: those are buckets in which a server had any collector write a row, and a body spanning 8–10 s straddles several buckets. At cadence 109.7 s across 42 servers you expect ~5.7 servers touching an average bucket, clustering higher. No defect here.

What actually reaches 60 s

With cadence = 42 × R / C + 7.5:

lever use1 use2
current (R, C=4) 109.7 s 69.8 s
C = 5 89.3 s 57.3 s ✅
C = 8 58.6 s ✅ 38.6 s
R halved, C=4 (i.e. #2847) ~58 s ✅ —

Three independent levers each reach 60 s. The cheapest is the operator knob — it is a config change, not code.

Cost of raising C

Raising concurrency does not just compress existing work; it increases collection rate. use1 going 109.7 → 58.6 s is 1.87× more collections per unit time, so ~1.87× collector CPU. use1 currently sits at ~48% on 8 cores → ~90%, which is too tight. use2 at C=5 is only 1.22× → ~43%, comfortable.

Store pool: MaxPoolSize = 24. Post-#2822 each body borrows one store connection for its whole duration, so C=8 means 8 concurrent store connections against a pool shared with retention, alerting, observability and the MCP — still under 24, but the margin narrows and #2819 measured a 673 ms acquisition floor there.

Suggested sequencing

  1. procedure_stats is 9.2x slower on use1 than use2 for identical row counts, and is 69% of use1's collection body #2847 first — procedure_stats is ~4.9 s of use1's 9.73 s gate-held body, so fixing it roughly halves R and cuts CPU per collection.
  2. Then a modest C bump (4 → 5) rather than 4 → 8, which lands ~54 s at ~1.2× current CPU instead of ~1.9×.

Effect on open issues

Not done

max_concurrent_sweeps is a production config change with a real CPU cost, so it is Erik's call rather than mine. No code changed, nothing deployed, no box restarted.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions