Skip to content

Two cross-region servers sit at 100% of their sweep budget — collection body exceeds the 60s cadence every cycle, so they collect at half rate #2296

Description

@erikdarlingdata

Found during the 2026-08-16 dogfood watch on prod-sql-use2-monitor-01. This is the only recurring warning left on that box after the decommissioned registrations were removed: ~50/hour, steady across four hourly windows.

[prod-sql-use2-multi-01...] collection body has not completed after 62s of execution - skipping relaunch
[prod-sql-use2-multi-53...] collection body has not completed after 61s of execution - skipping relaunch

It is not a fault, and that is the point worth recording

get_collection_health on multi-01 reports all 40 collectors HEALTHY, 0 errors, 0 yields, 0% failure rate. Nothing is failing. The skip is back-pressure working exactly as designed — the previous body is still running, so relaunching would pile up.

The cause is arithmetic, visible in the per-collector averages:

procedure_stats   22,141 ms
query_store       16,590 ms
plan_correction   13,544 ms
query_stats        8,437 ms
                 ---------
                  ≈ 60.6 s   against a 60 s sweep

Four collectors consume the entire cadence on their own. So the body cannot finish inside one interval, every interval, and the relaunch is skipped every time.

The consequence

These two servers effectively collect at half the nominal cadence. Nothing reports that as degraded — every collector is HEALTHY and the log line reads like a transient. A reader looking at collection freshness sees "Warning" status at worst, and the health tool sees nothing wrong at all, because from each collector's own point of view nothing is.

Why these two specifically

Both registrations resolve to us-east-1 endpoints (prod-sql-use1-multi-01, prod-sql-use1-multi-53) while the collecting box is in us-east-2. Every round trip pays cross-region latency, which inflates all four of the expensive collectors at once. The other 40 servers on the same box, collected in-region, do not produce this warning.

Nine registrations on this box currently point at use1 endpoints; only these two saturate, so latency is necessary but not sufficient — these are presumably also the busiest of the nine.

Options, roughly in order of appeal

  1. Move them to the us-east-1 box. prod-sql-use1-monitor-01 exists, is built, and is in-region for them. It is currently blocked on the darling_monitor role, so this is gated on that credential rather than on any code.
  2. Give the sweep a longer interval for saturated servers, so the body fits. Changes the data grain for those servers.
  3. Make the saturation visible rather than fixing it — surface "body exceeded its cadence" as a health-tool signal instead of only a log line, so half-rate collection is not something you have to notice by reading warnings. Worth doing regardless of which of 1 or 2 happens, because today a server can silently halve its own resolution while every collector reports HEALTHY.

Measurement notes

Counts are stable, not rising: 55/hr, 50/hr, 50/hr, 53/hr across the four windows since the decommissioned servers were removed at 14:05 UTC. Durations above are the avg_duration_ms from get_collection_health on multi-01 at 15:28 UTC.

Activity

  1. erikdarlingdata commented on Aug 17, 2026

    @erikdarlingdata
    OwnerAuthor

    Taking option 3 (make the saturation visible), which the issue already marks worth doing regardless of 1 vs 2 — mapping the seams now, PR to follow. Options 1 and 2 stay open: 1 is gated on the use1 darling_monitor credential (that box's binaries are current as of today's nightly, so it's install-ready the moment the role exists), and 2 changes the data grain for the saturated servers, which is Erik's call.

  2. erikdarlingdata commented on Aug 17, 2026

    @erikdarlingdata
    OwnerAuthor

    Correction from Erik that unblocks option 1: the use1 box can use the SAME credentials as use2 for SQL Server targets — the darling_monitor role gap applies only to the Aurora PostgreSQL clusters. So moving prod-sql-use1-multi-01/53 in-region is actionable now: start the use1 Darling service with those two registrations (its binaries are already on today's nightly), verify in-region collection, then remove them from use2's registry. Queued behind the #2295 backfill, which has the disk clock. Option 3's visibility PR (#2308) is review-clean and merging as CI recovers from today's GitHub outage.

  3. erikdarlingdata commented on Aug 17, 2026

    @erikdarlingdata
    OwnerAuthor

    Option 1 executed — the two saturated servers now collect in-region. On prod-sql-use1-monitor-01: the service is installed and Running (fresh store at v74, 49/49 hypertables, 17/17 retention policies armed from birth), the shared SQL credentials passed --test-connection on both targets, and after the first minutes of collection both servers show all-SUCCESS runs (multi-01: 59 runs / 134k rows; multi-53: 50 runs / 27k rows; 0 ERROR lines) at 1–3 ms in-region latency instead of cross-region round trips. Both registrations are removed from use2's registry (history retained), so its sweep no longer carries the two bodies that couldn't fit their cadence — the ~50/hour skip-relaunch warnings should stop within a sweep; I'll confirm on the next log check. The use1 box's original Aurora-placeholder config is preserved as darling.aurora-pending.json; note the Aurora enrollment now needs to go through the store (add_servers/viewer), not the file, since the store has seeded.

    With option 1 done, option 3's sweep_pressure verdict merged (#2308), and option 2 unnecessary, this issue closes once the use2 warning stream is confirmed quiet.

  4. erikdarlingdata commented on Aug 17, 2026

    @erikdarlingdata
    OwnerAuthor

    Resolved, and the fix grew into a fleet re-architecture once Erik corrected the topology model: use1 instances are writable PRIMARIES, use2 instances are their cross-region readable replicas, and both sides of every pair need monitoring. The RDS replication map shows 42 symmetric pairs — and only 9 primaries were monitored anywhere (7 cross-region from the use2 box, which is what this issue's saturation was; the other 33 were dark).

    End state, built and verified today:

    • prod-sql-use1-monitor-01 now monitors all 42 primaries in-region (fresh store at v74, retention armed from birth; all 42 onboarded via add_servers with per-server connection tests, 0 failures; in-region latency 1–3 ms).
    • prod-sql-use2-monitor-01 now monitors all 42 replicas in-region — the 9 dark replicas were registered, and the 7 cross-region primary registrations removed after their use1 coverage was verified collecting (history retained under the old identities).
    • The sweep effect was immediate: the ~50/hour skip-relaunch warnings stopped at the removals (0 since), and use2's fleet went from chronically half-Warning to 42/42 Online — the Warning haze was the cross-region bodies stretching every sweep.

    One mechanism note for the record: add_servers resolves duplicates against display names, and the old cross-region rows carried the replica hostnames AS display names — so adding a replica while its pair's old row still existed merged into that row instead of creating a new one, and the subsequent removal took both out. Remove-then-add is the safe order for a re-target; the 7 affected replicas were re-added cleanly within minutes. Possibly worth a guardrail issue: an add that MERGES into an existing row reports itself as "added", which reads as a new registration when it isn't.

    The sweep_pressure verdict from #2308 now stands watch on both boxes for any recurrence.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions