Repository navigation
Treat a stale running_jobs snapshot as no evidence (#1812) - #1814
Merged
Merged
Conversation
All three editions read the latest running_jobs snapshot with no freshness guard; the engine's per-run cooldown key deliberately expires each pass, so a stale snapshot re-fired the same historical run every cooldown, forever. - Shared contract carries evidence quality: AnomalousJobsResult with SnapshotIsFresh; adapters probe MAX(collection_time) first and skip the row read entirely when stale - The engine skips BOTH branches on stale evidence: no fire, no fabricated "jobs cleared", active state untouched - fresh evidence resumes real evaluation - Freshness bound = 3x the server's EFFECTIVE running_jobs cadence (floor 10 min), supplied by Lite's ScheduleManager and Darling's StoreConfigProvider resolution (live overrides) - Dashboard's SQL-side read gets the same bound in its house idiom (fixed 10-minute DATEADD) Verified: engine harness walks the exact reported sequence; Lite DuckDB fixture (stale/fresh/cadence-widened/empty); Darling live leg ages the seeded snapshot in place + cadence hook; bound watched red by mutation; Lite 1605/1605, Darling fast 3565, all builds zero warnings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
erikdarlingdata
added a commit
that referenced
this pull request
Jul 29, 2026
…key UPDATE The #1814 staleness test aged its snapshot by UPDATEing collection_time to now-2h. collection_time is the hypertable's partition key and TimescaleDB never re-routes a row on UPDATE, so between 00:00 and 02:00 UTC the new value crosses the midnight chunk boundary and violates the chunk's slice CHECK constraint (23514) - a nightly two-hour time bomb that passed its own CI and two nightlies because they all ran outside the window. First detonation was this release PR's 00:42 UTC run. A DELETE + re-INSERT routes to the correct chunk at any hour; the stale and relaxed-profile assertions are unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #1812.
The reported bug — confirmed in all three editions
Every edition read the LATEST running_jobs snapshot with no freshness guard, exactly as the reporter quoted. And the repetition at cooldown intervals had a second cause invisible from the query: the per-run cooldown key (
{server}:{job}:{start}) is deliberately dropped once it ages past the cooldown, so a stale snapshot whose rows keep matching re-arms and re-fires the same historical run every cooldown, forever.The fix — stale is NO evidence, in both directions
GetAnomalousJobsAsyncreturnsAnomalousJobsResult (SnapshotIsFresh, Jobs). Lite and Darling probeMAX(collection_time)first and return the stale result — the row read skipped entirely — when the newest snapshot is older than the bound (or absent).StoreConfigProvider.ResolveSchedulereading the live overrides field, so a control-plane reload reaches the very next check. Both fall back to the shared default.DATEADD(MINUTE, -10, SYSDATETIME()), matching its Agent-driven cadence and the file's existing staleness bounds). One documented behavioral delta: its legacy loop keeps its simpler shape, so a freshly-dead collector may emit one final "cleared" there where Lite/Darling now stay silent.Test plan
MaxSnapshotAge→ infinite fails the stale test)🤖 Generated with Claude Code