Repository navigation
CI flake: PgTarget anomaly/blocking worst-tile assertions fail on first attempts unrelated to the change #4274
Description
Activity
Cause confirmed and fixed in #fix/4274-pgtarget-tile-flake (PR to follow): both tests derived their
analysis window from a rawDateTime.UtcNowminute.AnomalyGate.EvaluateTilesskips any tile under
AnomalyThresholds.MinTileSamples(3), so when the wall clock put the seeded spike/chain's own tile at
1-2 samples, that tile was excluded from scoring and a different tile "won" worst-scoring.- Anomaly test: reproduced at a controlled anchor (spike tile = 2 samples) —
current_ms_per_sec
expected 3200, got 1500 (exact match to the CI failure). - Blocking test: reproduced at a controlled anchor (leading tile = 11 samples) —
avg_blocked_sessions
expected in [2.8, 3.0], got 2.7272727272727271 (exact match to the CI failure, including the digits).
Both tests are now fixed by anchoring the window end to the hour (
TruncateToHour(now).AddMinutes(-1),
always :59 of the prior hour) instead of the raw current minute, so the tile holding the seed is never
undersized regardless of wall time. Verified at 7 fixed clocks spread across an hour plus one crossing
UTC midnight, and the real full Darling suite (14070 total, 0 failed).One data point worth a coordinator look, not acted on here: at the bad anchor that reproduced the
anomaly-test flake, the fact'scurrent_ms_per_sec(1500) genuinely undersold the real spike, while
window_peak(3200) carried the correct value the whole time. That's #3653 A8 option B's tiling working
as designed — a spike that lands in a too-small edge tile of the SCORED window is real product behavior,
not a test bug — so I leftAnomalyGate/MinTileSamplesuntouched and only fixed the tests' own
non-determinism. Flagging in case the coordinator wants a follow-up on whether an edge tile that thin
should ever be allowed to out-score a full tile holding the same spike, rather than folding into the
adjacent tile or similar.- Anomaly test: reproduced at a controlled anchor (spike tile = 2 samples) —
Answering the question in issuecomment-5836299643: no product change. The worst-tile design (#3653 A8 option B) skips a tile with fewer than
MinTileSamplessamples on purpose, so a 1- or 2-sample tile can't decide the verdict.window_peakstill carries the real peak, as FX2 measured (3200). On the newest tile, the cost is a detection delay of at mostMinTileSamples - 1samples. PR #4323 fixes the tests' anchors. The two other raw anchors (PgTargetBlockingTestszero-history e2e,PgTargetAnomalyTeststhirty-day spike e2e) don't need the change: each window sits entirely inside a uniform series, so every full tile scores the same whatever the minute.🤖 Generated with Claude Code
- added a commit that references this issue
on Sep 25, 2026
What's failing
Two live PostgreSQL tests in the
PgTarget*anomaly/blocking family failed on first attempts, in Build runs whose change did not touch either test:fix/4226-viewer-timer-reads(Overview/ServerSummary reads)PgTargetAnomalyTests.TheAuroraWaitProfile_OneHotCollectionStaysQuiet_ASustainedShiftFires_AgainstDevPostgrescurrent_ms_per_secexpected 3200, got 1500#4259get_query_heatmap defaults#3953query-store interval fixPgTargetBlockingTests.ThirtyOneDaysOfLightBlocking_ThenAFourHourChainUnderAnIdleHead_YieldsTheChainStory_TheEvents_TheRunner_AndTheAnomalyInOneIncidentAssert.InRange(anomaly.Metadata["avg_blocked_sessions"], 2.8, 3.0)— actual 2.7272727272727271None of the three changes touch anomaly detection, tile scoring, or blocking analysis.
PgTargetAnomalyTestsrecurred on two unrelated changes (a PR and a later dev push), so it is a repeat, not a one-off.Why this is one issue, not two
Both failing assertions read a value that the code (per its own comments) derives from "the worst-scoring TILE" under
#3653 A8 option B— a target-local-hour bucket picked out of a multi-hour window, not the whole window's own mean/peak:PgTargetAnomalyTests:fact.Metadata["current_ms_per_sec"]should be the seeded spike (3,200) but read 1,500 — the flat background rate the rest of the window carries. That is consistent with the detector scoring a different hour-tile than the one holding the spike.PgTargetBlockingTests: the comment atPgTargetBlockingTests.cs:790-793already documents thatavg_blocked_sessions"comes from the WORST-SCORING TILE's own mean," pinned to2.8-3.0; the actual (2.727...) reads like a different tile's mean.Both tests build their windows from
DateTime.UtcNowat run time (TruncateToMinutes(DateTime.UtcNow),windowStart = end.AddHours(-4)), not a frozen clock, so if tile boundaries are wall-clock-hour-aligned rather than window-relative, which hour is "worst" — and therefore which value the fact reports — can depend on wherenowfalls relative to an hour boundary when the test runs. That would explain intermittent failure without shared-store pollution: both server IDs and their rows are test-owned and cleaned up (DeleteWaitGateRowsAsync,DeleteRowsAsync), so this isn't the job-history-style fleet-wide-read pollution #4235's own flake (fixed in #) turned out to have.This is a hypothesis from the failure shape and the tests' own comments, not a confirmed root cause — I have not traced
PgTargetAnomalyDetector's tile-selection code. That tracing, plus deciding whether the fix belongs in the detector or in these two tests, is real investigation, not a small fix, so it did not fit in the #4235 CI-unblock lane.Census method
gh run list --workflow Build --limit 80, each run's job list via the RESTactions/runs/<id>/jobsendpoint, and a grep of[FAIL]/Total:fromgh run view --job <id> --log-failedfor every job whosebuildstep succeeded (ruling out the PR's own compile break) and whose failing test's area the PR/push's own diff didn't touch, across Build runs from 2026-09-25 05:13Z through 09:41Z.