Skip to content

Treat a stale running_jobs snapshot as no evidence (#1812) - #1814

Merged
erikdarlingdata merged 1 commit into
devfrom
feature/1812-stale-jobs-snapshot
Jul 28, 2026
Merged

erikdarlingdata merged 1 commit into
devfrom
feature/1812-stale-jobs-snapshot

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Closes #1812.

The reported bug — confirmed in all three editions

Every edition read the LATEST running_jobs snapshot with no freshness guard, exactly as the reporter quoted. And the repetition at cooldown intervals had a second cause invisible from the query: the per-run cooldown key ({server}:{job}:{start}) is deliberately dropped once it ages past the cooldown, so a stale snapshot whose rows keep matching re-arms and re-fires the same historical run every cooldown, forever.

The fix — stale is NO evidence, in both directions

  • Shared contract carries evidence quality: GetAnomalousJobsAsync returns AnomalousJobsResult (SnapshotIsFresh, Jobs). Lite and Darling probe MAX(collection_time) first and return the stale result — the row read skipped entirely — when the newest snapshot is older than the bound (or absent).
  • The engine skips BOTH branches on stale evidence: no fire, and no fabricated "Long-Running Jobs Cleared" off a collector that merely stopped reporting. Active state stays untouched, so returning fresh evidence resumes real evaluation and fires or resolves from actual facts.
  • The bound derives from each server's EFFECTIVE cadence (3 missed cycles, floored at 10 minutes): the shipped profiles span 2–30 minutes, so a fixed bound is either blind for the fast profile or a gag for the slow one. Lite supplies its ScheduleManager lookup (the failed-jobs fetcher's hash-match idiom); Darling supplies StoreConfigProvider.ResolveSchedule reading the live overrides field, so a control-plane reload reaches the very next check. Both fall back to the shared default.
  • Dashboard parity: its own SQL-side read gets the same protection in its house idiom (fixed 10-minute DATEADD(MINUTE, -10, SYSDATETIME()), matching its Agent-driven cadence and the file's existing staleness bounds). One documented behavioral delta: its legacy loop keeps its simpler shape, so a freshly-dead collector may emit one final "cleared" there where Lite/Darling now stay silent.

Test plan

  • Engine harness walks the exact reported sequence: live fire → collector dies → two cooldowns pass with NO re-fire and NO fabricated resolution → fresh-and-empty evidence returns → the REAL resolution fires
  • Lite adapter against a real DuckDB store: stale → no evidence with rows skipped; fresh → rows; 60-minute cadence makes a 2-hour-old snapshot CURRENT (the hook genuinely widens the bound); empty store → no evidence
  • Darling live leg extended: ages the seeded snapshot in place past the bound → no evidence; the cadence hook proven live
  • Bound watched RED by mutation (MaxSnapshotAge → infinite fails the stale test)
  • Full suites: Lite 1605/1605, Darling fast 3565; Lite + Darling service/viewer + Dashboard builds all zero warnings
  • darling-pg CI leg

🤖 Generated with Claude Code

All three editions read the latest running_jobs snapshot with no
freshness guard; the engine's per-run cooldown key deliberately expires
each pass, so a stale snapshot re-fired the same historical run every
cooldown, forever.

- Shared contract carries evidence quality: AnomalousJobsResult with
  SnapshotIsFresh; adapters probe MAX(collection_time) first and skip
  the row read entirely when stale
- The engine skips BOTH branches on stale evidence: no fire, no
  fabricated "jobs cleared", active state untouched - fresh evidence
  resumes real evaluation
- Freshness bound = 3x the server's EFFECTIVE running_jobs cadence
  (floor 10 min), supplied by Lite's ScheduleManager and Darling's
  StoreConfigProvider resolution (live overrides)
- Dashboard's SQL-side read gets the same bound in its house idiom
  (fixed 10-minute DATEADD)

Verified: engine harness walks the exact reported sequence; Lite DuckDB
fixture (stale/fresh/cadence-widened/empty); Darling live leg ages the
seeded snapshot in place + cadence hook; bound watched red by mutation;
Lite 1605/1605, Darling fast 3565, all builds zero warnings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@erikdarlingdata
erikdarlingdata merged commit e205777 into dev Jul 28, 2026
4 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/1812-stale-jobs-snapshot branch July 28, 2026 19:16
erikdarlingdata added a commit that referenced this pull request Jul 29, 2026
…key UPDATE

The #1814 staleness test aged its snapshot by UPDATEing collection_time to
now-2h. collection_time is the hypertable's partition key and TimescaleDB
never re-routes a row on UPDATE, so between 00:00 and 02:00 UTC the new value
crosses the midnight chunk boundary and violates the chunk's slice CHECK
constraint (23514) - a nightly two-hour time bomb that passed its own CI and
two nightlies because they all ran outside the window. First detonation was
this release PR's 00:42 UTC run. A DELETE + re-INSERT routes to the correct
chunk at any hour; the stale and relaxed-profile assertions are unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant