You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Collection cadence has no fixed component: cadence = N x gate_held_time / max_concurrent_sweeps + tick/2 #2849
The "~55 s fixed component" in #2841does not exist. Cadence is fully explained by a purely multiplicative model derived from the launch loop, with no intercept beyond half a sweep tick. Measured on both boxes after today's resize, the model fits use1 to 0.4%.
The fire-and-track loop (#1553) launches every server every 15 s tick unless its previous body is still in flight. The concurrency gate is acquired inside the body (await gate.WaitAsync(stoppingToken) is the first statement of ProcessServerSweepAsync), so all 42 bodies launch and queue on the semaphore. It is a queue, not a round-robin.
Steady state therefore gives:
cadence ≈ (N × gate_held_time) / C + tick/2
where N = enabled servers, C = max_concurrent_sweeps, tick = s_sweepInterval (15 s). The tick/2 is the average wait for the next launch tick after a body completes.
Measurements (window >= 2026-09-03 14:00Z, both boxes m7i.2xlarge, post-resize)
use1
use2
config_service.max_concurrent_sweeps
4
4
enabled servers
42
42 (61 registered, 19 disabled)
body span p50 / mean (collection_log first→last row per sweep)
use1 fits almost exactly. use2's 2.13 s residual is gate-held time that produces no collection_log row and so falls outside a span measured between first and last row — consistent with the honest-empties convention (a gated-off collector gets no row at all, and use2's Query Store is dead by design since #2296).
It fitted a straight line through two points whose x-values were synthetic sums of the 1-minute tier (7,061 ms and 2,200 ms), not measured gate-held time. Measured gate-held time is 9.73 s and 5.93 s — ratio 1.64, versus the synthetic ratio of 3.21. Fitting a line through mis-scaled x-values manufactures an intercept. There is no fixed term.
max_concurrent_sweeps reconciliation
The knob is at its documented default of 4 on both stores — it is not set high, and the limit is being enforced. The earlier "10 and 23 distinct servers per 15 s bucket" observation is not evidence of concurrency above 4: those are buckets in which a server had any collector write a row, and a body spanning 8–10 s straddles several buckets. At cadence 109.7 s across 42 servers you expect ~5.7 servers touching an average bucket, clustering higher. No defect here.
Three independent levers each reach 60 s. The cheapest is the operator knob — it is a config change, not code.
Cost of raising C
Raising concurrency does not just compress existing work; it increases collection rate. use1 going 109.7 → 58.6 s is 1.87× more collections per unit time, so ~1.87× collector CPU. use1 currently sits at ~48% on 8 cores → ~90%, which is too tight. use2 at C=5 is only 1.22× → ~43%, comfortable.
Store pool: MaxPoolSize = 24. Post-#2822 each body borrows one store connection for its whole duration, so C=8 means 8 concurrent store connections against a pool shared with retention, alerting, observability and the MCP — still under 24, but the margin narrows and #2819 measured a 673 ms acquisition floor there.
max_concurrent_sweeps is a production config change with a real CPU cost, so it is Erik's call rather than mine. No code changed, nothing deployed, no box restarted.
Summary
The "~55 s fixed component" in #2841 does not exist. Cadence is fully explained by a purely multiplicative model derived from the launch loop, with no intercept beyond half a sweep tick. Measured on both boxes after today's resize, the model fits use1 to 0.4%.
This reframes #2841 and promotes #2847.
The mechanism (from
DarlingWorkerlaunch loop)The fire-and-track loop (#1553) launches every server every 15 s tick unless its previous body is still in flight. The concurrency gate is acquired inside the body (
await gate.WaitAsync(stoppingToken)is the first statement ofProcessServerSweepAsync), so all 42 bodies launch and queue on the semaphore. It is a queue, not a round-robin.Steady state therefore gives:
where
N= enabled servers,C=max_concurrent_sweeps,tick=s_sweepInterval(15 s). Thetick/2is the average wait for the next launch tick after a body completes.Measurements (window >= 2026-09-03 14:00Z, both boxes m7i.2xlarge, post-resize)
config_service.max_concurrent_sweepscollection_logfirst→last row per sweep)Model fit
Solving for gate-held time from observed cadence:
R= (cadence − 7.5) × C / Nuse1 fits almost exactly. use2's 2.13 s residual is gate-held time that produces no
collection_logrow and so falls outside a span measured between first and last row — consistent with the honest-empties convention (a gated-off collector gets no row at all, and use2's Query Store is dead by design since #2296).Why #2841 saw a 55 s intercept
It fitted a straight line through two points whose x-values were synthetic sums of the 1-minute tier (7,061 ms and 2,200 ms), not measured gate-held time. Measured gate-held time is 9.73 s and 5.93 s — ratio 1.64, versus the synthetic ratio of 3.21. Fitting a line through mis-scaled x-values manufactures an intercept. There is no fixed term.
max_concurrent_sweepsreconciliationThe knob is at its documented default of 4 on both stores — it is not set high, and the limit is being enforced. The earlier "10 and 23 distinct servers per 15 s bucket" observation is not evidence of concurrency above 4: those are buckets in which a server had any collector write a row, and a body spanning 8–10 s straddles several buckets. At cadence 109.7 s across 42 servers you expect ~5.7 servers touching an average bucket, clustering higher. No defect here.
What actually reaches 60 s
With
cadence = 42 × R / C + 7.5:R, C=4)Rhalved, C=4 (i.e. #2847)Three independent levers each reach 60 s. The cheapest is the operator knob — it is a config change, not code.
Cost of raising
CRaising concurrency does not just compress existing work; it increases collection rate. use1 going 109.7 → 58.6 s is 1.87× more collections per unit time, so ~1.87× collector CPU. use1 currently sits at ~48% on 8 cores → ~90%, which is too tight. use2 at C=5 is only 1.22× → ~43%, comfortable.
Store pool:
MaxPoolSize = 24. Post-#2822 each body borrows one store connection for its whole duration, so C=8 means 8 concurrent store connections against a pool shared with retention, alerting, observability and the MCP — still under 24, but the margin narrows and #2819 measured a 673 ms acquisition floor there.Suggested sequencing
procedure_statsis ~4.9 s of use1's 9.73 s gate-held body, so fixing it roughly halvesRand cuts CPU per collection.Cbump (4 → 5) rather than 4 → 8, which lands ~54 s at ~1.2× current CPU instead of ~1.9×.Effect on open issues
cadence ∝ body. But its conclusion that ~1.5 min is an unreachable floor is wrong — there is no fixed term, and both body reduction and concurrency reach 60 s.Not done
max_concurrent_sweepsis a production config change with a real CPU cost, so it is Erik's call rather than mine. No code changed, nothing deployed, no box restarted.