Repository navigation
CI: distinct-chunk counting for job_history's two-scan floor proof (flake from #4235) - #4275
Merged
Merged
Conversation
…of (flake from #4235) EventWindowedReadsAreBoundedLivePostgresTests failed dev's Build runs 36115696476 and 36117118128 on job_history's AssertChunksAsync. #4229's GROUP BY fix split the floored job_history read into two independently floor-bounded scans (job_stats CTE plus the base join), so the floored plan now visits every in-window chunk twice. PlanChunkScans.Count counts scan nodes, not physical chunks, so that doubling ate the two-chunk reduction margin the test asserted, even on a freshly created database with nothing else in it (measured old=7, new=6 raw nodes, reduction 1). Add PlanChunkScans.DistinctChunkCount and use it in AssertChunksAsync so a chunk scanned twice by two Append branches over the same hypertable counts once. Measured after the fix: old=7, new=3 distinct chunks, reduction 4, comfortably clearing minReduction: 2. Confirmed the assertion still catches a real regression: with the floor neutralized for one read, old=7, new=7, and the test correctly fails. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TszxYhJJbTEh4LrZ56NYo3
erikdarlingdata
marked this pull request as ready for review
September 25, 2026 10:34
This was referenced Sep 25, 2026
3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the
EventWindowedReadsAreBoundedLivePostgresTestsflake blocking every armed PR's required check.Census
Workflow:
gh run list --workflow Build --limit 80, covering 2026-09-25 05:13Z-09:41Z (the full window CI's volume returned). Jobs per run: REST endpointactions/runs/<id>/jobs. Failure output:gh run view --job <id> --log-failed, filtered to[FAIL]andTotal:for every failed PG-tests job.Excluded as own-PR compile breaks (
buildalso failed). Runs:36098645834,36099235605,36100129482,36100816723,36104266862,36105160075,36106541517,36110524470,36112451005,36112955353,36114913004,36117448144,36119112979.Two more excluded (own bug, not flake).
36115560721:fix/4198-collection-health-default, Lite'sMcpToolsListBudgetTestsceiling broken by that PR's own description change.36097633466:fix/4226-viewer-timer-reads, failed onViewerW2aLivePostgresTests, the read that PR was changing.Cross-cutting failures, where the change did not touch the failing area:
36109281530#3953query-store interval fixPgTargetBlockingTests.ThirtyOneDaysOfLightBlocking...avg_blocked_sessionsInRange(2.8,3.0), actual 2.72736100667038fix/4226-viewer-timer-reads(Overview reads)PgTargetAnomalyTests.TheAuroraWaitProfile...current_ms_per_secexpected 3200, actual 150036115696476#4259get_query_heatmap defaultsEventWindowedReadsAreBoundedLivePostgresTests(job_history) +PgTargetAnomalyTests.TheAuroraWaitProfile...(same as above)36117118128#4228AG readsEventWindowedReadsAreBoundedLivePostgresTests(job_history) onlyBaseline confirmed green:
36111242422(08:07Z) and36112268301(08:19Z), both fully passing, both after#4235merged (06:11Z) and before the two failing runs.That is two flake families, not one.
PgTargetAnomalyTests/PgTargetBlockingTestsshare a "worst-scoring tile" pattern (#3653 A8 option B), unrelated to chunk counting or job history. I filed #4274 for that one with its own evidence rather than fixing it here. TracingPgTargetAnomalyDetector's tile selection is real investigation, not a small fix.Root cause (confirmed with numbers, not the shared-pollution theory)
The working theory going in was shared-store pollution shifting chunk counts based on CI's shard assignment. That is not what is happening here. I reproduced the job_history failure alone, on a freshly created
darlingtest, with zero other classes run first:That is the same reduction (1) the polluted CI runs showed (old=8, new=7 there). The +1 on both sides in CI is a
servers-table chunk unrelated to job_history, so it nets zero effect on the difference. The real cause is structural, not pollution:ChunkIntervalDays = 1(TimescaleSupport.cs). Job_history's chunks are calendar-day-aligned.EventWindowFloor.Forsets the floor to exactlyWindowStart - 1 day, which isnow - 2 days. That is an exact instant, never a chunk boundary.#4229's own GROUP BY fix splitBuildJobHistorySql's single scan into two independently floor-bounded scans, ajob_statsCTE plus the base join. So the floored plan visits every in-window chunk twice.PlanChunkScans.Countcounts scan-node lines, not physical chunks. A 48-hour lookback from an arbitrary instant almost never fits in 2 calendar days, so 3 in-window calendar-day chunks get scanned, each twice.newreads as 6 nodes.old's single scan across all 7 existing chunks reads as 7. Reduction by node count: 1. The test's ownminReduction: 2was already discounted from the default 3 for this doubling, but it is still unreachable by construction. No amount of database cleanup changes that math.Fix
Darling/Darling.Tests/PlanChunkScans.cs: addedDistinctChunkCount, counting unique_hyper_N_M_chunknames instead of scan-node occurrences.Countand its existing tests (PlanChunkScansTests.cs) are untouched, so other callers keep today's node-counting behavior.Darling/Darling.Tests/EventWindowedReadsAreBoundedLivePostgresTests.cs:AssertChunksAsyncnow comparesDistinctChunkCount. A chunk visited by two scan nodes over the same hypertable now counts once, which is what "the floor prunes out-of-window chunks" actually means. The other fiveAssertChunksAsynccalls each scan two different hypertables. Their chunk names never collide, so this change does not affect them.Measured after the fix, fresh database: old=7, new=3 distinct chunks, reduction 4 (comfortable margin over
minReduction: 2, tolerant of a modest pollution shift either side).Proof it still catches a real regression
Per the lane orders, deleted the floor from one read to prove the assertion still fails without it. Swapped the job_history call's floor parameter for a no-op (
1970-01-01, changing nothing else) and reran:Correctly fails (reduction 0). Reverted before committing, so the pushed diff only touches the two files above.
Test plan
dotnet build Darling/Darling.Tests/Darling.Tests.csproj -c Debug: 0 Warning(s), 0 Error(s)EventWindowedReadsAreBoundedLivePostgresTestsalone, freshdarlingtest, 5 consecutive runs: 5/5 pass (Total: 2, Failed: 0each time)Darling.Testssuite, once, freshdarlingtest:Total: 13893, Errors: 0, Failed: 1, Skipped: 47, Not Run: 1, Time: 552.377s. The one failure,ServerListAndSummaryPlanShapeTests.TheShippedReads_TouchFarFewerChunks..., did not appear in any CI run in the census. It passed alone on a fresh database: an artifact of running all 13,893 tests serially. A single CI shard runs about a third of the classes, so it accumulates far lesscollection_logchunk state than an unsharded run. Per the lane orders' pre-existing-failure rule, a failure that clears on a fresh database alone is not filed.Issues filed
PgTargetAnomalyTests/PgTargetBlockingTestsworst-tile flake family (2+1 occurrences in this census), root cause not confirmed, needs real tracing into tile selection.For the coordinator to double-check
DistinctChunkCountreasoning holds for job_history (same hypertable, two scans). I did not audit every other live plan-shape test for the same double-scan-over-one-hypertable shape. Any such test has the identical latent bug.#4198's in-flight PRs or thefix/4198-*build failures. Those are each branch's own compile breaks, unrelated to this flake.