Skip to content

Render the sweep's memory: the Fleet Sweeps page (#3466, lane 3 of 4) - #3472

Merged
erikdarlingdata merged 2 commits into
devfrom
feat/3466-sweep-web-feed
Sep 16, 2026
Merged

erikdarlingdata merged 2 commits into
devfrom
feat/3466-sweep-web-feed

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Render the sweep's memory: the web feed at configurable spans, where mute costs and dead instruments stay visible

Lane 3 of #3466's approved four-lane build order, following lanes 1–2, both now merged (#3470 V123, #3471 V124 — including the sighting-semantics follow-up this page's "Last seen" column and its last_seen_at DESC ordering depend on: the stamp moves only on sightings, so the column means what it says). What lane 1 persisted and lane 2 writes on cadence, this lane makes VISIBLE: the read-only /api/sweeps surface on the Darling web dashboard and the Fleet Sweeps page that renders it — the per-sweep stateful workhorse the owner's delivery ruling assigned to the web ("the web feed for it should be configurable spans of time though, default hourly"), with channel delivery deliberately absent because it is lane 4's daily-ceiling rollup. This is also the lane the feature becomes user-visible in, so the changelog entry lanes 1–2 deliberately deferred lands here.

The endpoints: dedicated reads, like /api/fleet, not tool mirrors

Four GET routes under /api/sweeps, mapped from MapAll beside the triage endpoint and shaped by the /api/fleet//api/ag precedent — dedicated DTO endpoints, not /api/read/{tool} mirrors, because the sweep has no read tool yet: lane 4's get_sweep_reports arrives against the SAME presentation builders these routes serve. ExcludedToolNames is therefore untouched and the tool-catalog parity pin has nothing of this surface's to count, verified rather than assumed.

  • GET /api/sweeps — the timeline: runs inside the caller's span, newest first, each with its document embedded. hours defaults to 1 (the ruling's default-hourly view; at the shipped hourly cadence that lands on the newest sweep with the span control as the way into history) and both knobs go through McpHelpers.ValidateWindow — the SAME authority every windowed MCP read applies, so the two surfaces cannot disagree about what a bad span or a future as_of anchor means, and the refusal prose is one vocabulary. An out-of-range span is REFUSED, never clamped, and an unreadable one is refused rather than defaulted (the get_collection_log cannot answer the slow-run question: a 200-row newest-first cap makes hours_back irrelevant, with no collector or duration filter and no truncation disclosure #3287 filter discipline: a caller who mistyped a window must not receive a complete-looking answer to a different one). Both resolved bounds are echoed on the response — The analysis family cannot be anchored at a past window, and it is the one an incident starts with #2506's read end: a response that does not say what window it covers invites the caller to assume the one it asked for.
  • GET /api/sweeps/latest — the landing document. 404 with an honest sentence on a store with no sweep yet; a store FAULT is a 500, never a 404 that reads as "no sweeps".
  • GET /api/sweeps/{id} — the timeline click-through. 404 means genuinely absent (pruned, or never recorded); the store read for it throws on a fault for exactly that distinction.
  • GET /api/sweeps/watch-items — the worklist. The DEFAULT view is open ∪ carried (lane 1 review's ruling: "open right now" means the union, because the open state lasts exactly one sweep by design), served by ONE new store read rather than two calls at two instants; ?state= narrows to any of the machine's four states by exact spelling, and an unknown state is refused naming the legal values — an empty answer to a typo is indistinguishable from a clean worklist.

Seats: every route is a GET, so both seats reach all of it — nothing consults CanEdit, exactly per DarlingWebSeat.IsRequestAllowed's group-level gate (a sweep report is what the viewer seat exists to grant), and the auth middleware 401s a no-seat caller before any handler runs. The reads run on the host's least-privilege viewer pool, whose blanket collect SELECT already covers the sweep tables (lane 1's no-ACL rung, verified there).

The cadence knob is not touched. The page displays the effective cadence from /api/read/get_alert_settings' fleet_sweep group — the read every settings surface makes — and carries no write affordance at all; lane 2's V124, the Settings window and update_alert_settings own the writes. Pinned as a source census on the page (reads get_alert_settings, contains no apiSend, never names the write tool).

The shared presentation layer, so lane 4 cannot fork the wire

FleetSweepPresentation is the reshaping layer the lane brief demanded instead of per-surface forks: the store's reads stay row-shaped (the engine joins them relationally), and the RENDER reshaping happens once — the web routes serve these builders today, and lane 4's get_sweep_reports serves the same ones. What the reshaping actually is:

  • Documents ride EMBEDDED, not re-escaped: report_json, the liveness block and every evidence payload are parsed and embedded as objects (the BuildFullRuleNode definition-embedding precedent) — a string would make every client parse a document out of a document, and an agent on lane 4's tool would burn its window on escape characters. A payload that does not parse is carried VERBATIM as a string, failing toward visible rather than toward an absence that looks deliberate.
  • The honesty rules are restated on the wire: every run node carries alerts_enabled (the spec's on-every-sweep mute header) and instruments_alive unconditionally; the detail shape carries would_have_paged whenever the sweep ran under master-off INCLUDING EMPTY ("muted, and nothing would have paged" is a statement the operator is owed) and OMITS the key on an alerts-on sweep, where an empty array would claim a check the engine never made.
  • Ledger rows get server names JOINED from the same sweep's verdicts — a row keyed on a bare id would make the muted-mode contract's one table the one table an operator cannot read; a server the verdicts do not carry keeps its id and an honest null.
  • Verdicts carry band_changed pre-computed (the diff is the feature's acceptance test, so the wire states it) with the signals evidence embedded beside the band, verdict-beside-inputs.
  • Watch items carry the hysteresis BARS beside the counters (entry_bar_sweeps/exit_bar_sweeps, spliced from the state machine's constants): "1 miss" only means something beside "of 2 to close", and a client that hardcoded 2 would silently mis-render the day per-item thresholds arrive as data (the Retention Held's warn/critical ratios are compile-time constants, so the alert an operator most needs to tune is the one alert that cannot be #3297 route lane 1's docs reserve). The fleet-scope sentinel is named on the wire (fleet_scope) so no client has to know that server id 0 means the fleet.

The store gained exactly two reads, both pinned in lane 1's rung file

  • GetOpenAndCarriedWatchItemsSql / GetOpenAndCarriedWatchItemsAsync — the one-call open ∪ carried default view, both literals spliced from the state machine's constants (the active read's rule), pending and closed deliberately outside it, same ORDER BY as the single-state read so the two views cannot disagree about "newest first". Presentation posture: logs and degrades.
  • GetRunSql / GetSweepAsync — the by-id detail read, sharing RunColumns with the other two run reads and the one mapper. It THROWS on a store fault, deliberately unlike the span read beside it: this serves the detail route, whose degraded rendering is its own HTTP error body — a fault swallowed into null would be answered as a 404, telling an operator a sweep the store holds does not exist. Null means exactly absent.

Both join FleetSweepStateRungTests' statement census (no now(), parameterized, the derived state literals — quoted, because the bare word closed sits inside closed_sweep_id).

The page: the render carries the feature's contracts

js/pages/sweeps.js + a #/sweeps route and nav entry, in the existing idioms (module-level state surviving the 60s poll like the fleet page's sort, el()/textContent only per R4, pre-judged values rendered and never re-derived per R1). The layout: the sweep document (latest, or the clicked sweep), the timeline at the configured span, and the watch-item worklist.

  • The mute header renders on every muted sweep (an amber role="status" banner) and the would-have-paged ledger renders whenever the document carries one — including empty, with the sentence saying the mute cost nothing this sweep.
  • Quiet-is-not-clean applies to the render: instruments_alive: false is a red role="alert" banner on the document, a red edge on the document card, and a red edge + INSTRUMENTS DEAD badge on the timeline row — a dead-instrument sweep cannot be mistaken for a green one at any zoom.
  • The diff leads: band transitions with from/to and the stored reason, arrivals and departures, and a first sweep says "no previous sweep to diff against — this document states absolutes" rather than implying a quiet history.
  • The liveness block renders whole: the verdict, the pass-counter reading against the previous sweep's persisted baseline (advancing / frozen / no baseline), the restart-reset note, per-server read faults, and every judgeability note the engine wrote.
  • Watch items render their hysteresis position against the API's bars ("1 of 2 quiet to close"), with first/last-seen and the standing evidence behind a disclosure; a "Show closed" toggle asks for the history by name.
  • The per-server verdict bands are the daily-summary calculator's labels; its "No Data" is not a CSS-safe class and renders in the neutral Unknown treatment, stated in the one helper that maps it.

Censuses moved, tests, verification

  • DarlingWebOidcTests.IsRequestAllowed_Matrix gains the sweep rows: both seats GET everything, POST /api/sweeps born-gated.
  • DarlingWebAssetsTests pins js/pages/sweeps.js into the build-output census.
  • FleetSweepStateRungTests pins the two new statements (above).
  • FleetSweepWebFeedTests (new): the span refusal table (unreadable / zero / negative / over-max hours, bad and future anchors — through the shared MCP authority), the watch-state list == the machine's constants with wrong-case refused, the default-span-is-one-hour ruling, every presentation shape above (mute header and liveness always present, documents embedded, unparseable-carried-verbatim, ledger present-empty-under-mute and absent-under-alerts-on, names joined with honest nulls, band_changed truth table, bars and fleet-scope on the watch payload, window echo) — and the frontend parity pins in the AlertTrayStatusLoggedTests idiom, JS being untested by convention: SWEEP_WATCH_STATE_LABELS parsed strictly and compared key-by-key to the state machine's constants, the dead-instruments role="alert" branch and the ledger-whenever-carried branch located by source, the bars-not-hardcoded pin, the cadence-is-display-only pin, and the shell/router wiring pin.
  • ExcludedToolNames untouched; the /api/read parity pin unaffected (no new tool this lane). README's web-dashboard sentence and Darling/README's "What you see" blurb gain the Fleet Sweeps page; the CHANGELOG entry for the whole feature lands here per lanes 1–2's stated disposition (the lane that makes it visible owns the entry).

Development on macOS, where the test assemblies build but cannot run, so verification was: the full solution builds clean (0 errors; all warnings pre-existing in untouched files), and an out-of-band harness referencing the built Service assembly ran 53 checks — every pure binding/presentation behaviour above (the internal endpoint helpers reached by reflection), plus the live half against a real TimescaleDB pg17: full-ladder migration, a REAL FleetSweepEngine.RunAsync sweep plus a planted master-off sweep with watch items in all four states, the two new store reads over those rows (by-id round-trip and null-on-absent, open ∪ carried returning exactly the open and carried rows while the closed row stays reachable by name), and the routes served end-to-end by a real Kestrel host mapped through the real DarlingFleetSweepEndpoints.Map — the empty store's honest 404-latest and empty-window timeline, both sweeps on the wire with embedded documents, every 400 refusal in the shared prose, the mute header and name-joined ledger on latest, the ledger key absent on the alerts-on sweep, and the watch default/named/refused triple. All pass.

What lane 4 consumes from this PR

  • FleetSweepPresentation — get_sweep_reports serializes the SAME builders these routes serve (BuildTimelineNode / BuildSweepDetailNode / BuildWatchItemsNode), so the MCP surface and the web feed cannot drift; the embedded-documents rule is exactly what an agent consumer needs.
  • The store reads as shaped here — the daily rollup ranks over the same verdict/ledger rows; GetOpenAndCarriedWatchItemsAsync is the "open right now" read for the rollup's watch-transitions section, and GetSweepAsync the anchor for citing a specific sweep.
  • The wire vocabulary — would_have_paged present-iff-muted, band_changed, fleet_scope, and the bars-on-the-payload contract are now pinned shapes lane 4 inherits rather than decisions it re-makes.

Part of #3466 — depends on lane 1 (feat/3466-sweep-state-store, V123) and lane 2 (feat/3466-sweep-engine, V124); lane 4 (daily channel rollup ceiling + get_sweep_reports) follows against these shapes.

…mute costs and dead instruments stay visible (#3466)

Lane 3 of 4, stacked on the V123 state store and the V124 engine: the read-only
/api/sweeps surface (dedicated reads like /api/fleet, never tool mirrors — the
tool arrives in lane 4 against the same builders) and the Fleet Sweeps dashboard
page. The timeline serves caller-configurable spans, default one hour per the
owner's ruling, validated through the same McpHelpers.ValidateWindow authority
every windowed MCP read applies — refusing rather than clamping or defaulting,
and echoing the window it actually served. The watch worklist's default view is
open UNION carried in one new store read (open lasts exactly one sweep by
design), with any named state reachable by exact spelling and typos refused.

FleetSweepPresentation is the shared reshaping layer both presentation surfaces
serve: documents embedded as objects rather than re-escaped strings, ledger rows
carrying server names joined from the same sweep's verdicts, band transitions
pre-computed, and the hysteresis bars riding the payload beside the counters
they bound. The honesty rules are restated on the wire and in the render: the
mute header on every sweep with the would-have-paged ledger present whenever one
was taken under mute — empty included — and absent under alerts-on where an
empty array would claim a check never made; and a sweep that could not prove its
instruments renders LOUDLY, a red role=alert banner on the document and a marked
timeline row, because quiet-is-not-clean applies to CSS too. The cadence knob is
display-only here, read from get_alert_settings' fleet_sweep group.

Censuses moved: the seat matrix gains the sweep rows (both seats GET, unsafe
born-gated), the asset census gains pages/sweeps.js, the rung tests pin the two
new store statements, and FleetSweepWebFeedTests pins the span/state refusal
tables, every wire shape, and the frontend parity in the
AlertTrayStatusLoggedTests idiom (SWEEP_WATCH_STATE_LABELS parsed strictly
against the state machine's constants). ExcludedToolNames untouched. The
changelog entry lanes 1-2 deferred lands here, the lane that makes the feature
visible. Verified out-of-band on macOS: full solution builds clean, and a
53-check harness ran the pure half plus a live TimescaleDB pg17 pass — real
engine-written rows served end-to-end by a real Kestrel host through the real
Map, refusals, 404s, mute header, joined ledger and watch views all on the wire.
@erikdarlingdata
erikdarlingdata enabled auto-merge (squash) September 16, 2026 03:34
@claude

claude Bot commented Sep 16, 2026

Copy link
Copy Markdown

Reviewed the diff against dev (all 15 files). Summary: this is a clean, well-scoped addition — I didn't find correctness bugs, security issues, or Lite/Darling parity drift worth flagging as inline comments. Specifics checked:

Security

  • state (watch-item filter) is properly parameterized ($1) in GetWatchItemsByStateSql — not concatenated.
  • The only string-concatenated SQL literals (GetOpenAndCarriedWatchItemsSql's state IN (...)) are spliced from FleetSweepWatchStateMachine's hardcoded const string values, not caller input — no injection surface, matches the existing GetActiveWatchItemsSql precedent.
  • hours/as_of go through McpHelpers.ValidateWindow, the same authority MCP reads use; id is constrained via {id:long} at the route level.
  • JS renders exclusively through el()/textContent (confirmed in util.js's R4 rule) — no innerHTML path for any embedded document/evidence field, so the embedded report_json/evidence payloads can't become an XSS vector even though they're rendered as parsed objects.
  • Seat gating (DarlingWebSeat.IsRequestAllowed) is method-based (GET always allowed), so the new routes are correctly reachable by both seats and a stray POST /api/sweeps is refused by construction, without needing a route-specific check.

Correctness

  • Fault posture is consistent with the codebase's established split: GetSweepAsync/GetLatestSweepAsync throw (mapped to 500, never masquerading as 404); GetOpenAndCarriedWatchItemsAsync logs-and-degrades like its sibling presentation reads.
  • would_have_paged is correctly omitted (not empty-arrayed) when alerts_enabled: true, and present-but-possibly-empty under mute — verified both in FleetSweepPresentation and mirrored in sweeps.js's render logic.
  • Nullability of EmbedJson call sites lines up with each field's actual DB nullability (evidence_json/report_json/instrument_liveness_json are NOT NULL and called unguarded; verdict_json/watch-item evidence_json are nullable and guarded before the call).
  • Route registration order (/api/sweeps/latest, /api/sweeps/watch-items as literals vs /api/sweeps/{id:long}) is unambiguous under ASP.NET routing.

Parity

  • This is a Darling-only, fleet-wide feature (no Lite equivalent exists or is implied — Lite has no multi-server fleet concept), so the Lite/Darling parity rule doesn't apply here; confirmed no Lite files were touched and none should have been.
  • ExcludedToolNames/tool-catalog parity is correctly left untouched, as the PR description states (no MCP tool added in this lane).

Style

  • Matches CONTRIBUTING.md's Darling/PostgreSQL conventions (schema-qualified objects, naive UTC timestamps via .ToString("o") on UTC DateTimes, no T-SQL involved here since this is pure C#/PostgreSQL/JS).

Only a non-blocking nit: ValidateSpan's own "Invalid hours value" message and McpHelpers.ValidateWindow's downstream "Invalid hours_back value" message use different parameter names for the same knob — but the PR's own tests explicitly call this out and accept it as a known cosmetic inconsistency shared with the mute-rule endpoints, so I'm not raising it as an issue.

Nice work — the test coverage (FleetSweepWebFeedTests, the frontend source-scan pins) is thorough enough that I'd trust this to catch regressions on its own.

…ose naming the write tool it proves it never calls

Two censuses caught the page on its first Windows run: every rollup
figure must carry aria-describedby through the shared rollupTextId
derivation (#3031/#3045 - the tile closure now lives in the render
with a per-render id set, one construction path per page), and the
page's own reads-the-cadence-writes-nothing pin found the write
tool's name in the header comment. The prose now says what owns the
knob without spelling the tool it exists to not call.
@claude

claude Bot commented Sep 16, 2026

Copy link
Copy Markdown

Review

Went through the diff (FleetSweepStore, DarlingFleetSweepEndpoints, FleetSweepPresentation, sweeps.js, tests, css/wiring) against CONTRIBUTING.md.

Correctness — looks solid:

  • GetRunSql/GetSweepAsync reuse RunColumns/ReadRun; ordinals line up (verified against the existing ReadRun mapper).
  • GetOpenAndCarriedWatchItemsSql splices state literals from FleetSweepWatchStateMachine constants (not user input) and keeps the same ORDER BY as the single-state read, so the default view and a named-state view can't disagree about ordering.
  • Span validation (ValidateSpan) refuses unreadable/out-of-range hours and future as_of anchors rather than silently substituting a default — matches the #3287 discipline cited in the doc comments, and routes through the same McpHelpers.ValidateWindow the MCP tools use.
  • 404 vs 500 discipline is correct and consistent: GetSweepAsync/GetLatestSweepAsync throw on store faults (mapped to 500), while the child presentation reads (GetServerVerdictsAsync, GetWouldHavePagedAsync, GetOpenAndCarriedWatchItemsAsync) log-and-degrade to empty, so a child fault costs its section, not the whole page.
  • would_have_paged is correctly omitted (not empty-arrayed) when alerts_enabled is true, and present-including-empty when false — matches the honesty-rule described in the PR body, and is pinned by FleetSweepWebFeedTests.
  • ?state= is validated against KnownWatchStates before being handed to a parameterized query — no injection surface even though the state list happens to come from trusted constants.

Security — the new watch-item state filter and the sweep_id/hours/as_of query params are all parameterized or route-constrained ({id:long}); no string-built SQL. JS renders everything through el()'s textContent-only path (R4), including the JSON evidence blobs shown in <pre> — no innerHTML/XSS surface.

Lite/Darling parity — not applicable here: fleet sweeps are a Darling-only concept (multi-server fleet engine), and I confirmed there's no FleetSweep reference anywhere under Lite/, so this isn't a parity-drift case.

Style — Postgres-side SQL is parameterized, schema-qualified (collect.*), and the new endpoint file uses OPTION-equivalent discipline correctly N/A (no T-SQL here). C# style, XML doc conventions, and the two-store-parity rules in CONTRIBUTING.md don't apply since there's no Lite counterpart for this feature.

Minor observations (non-blocking):

  • DarlingFleetSweepEndpoints.ValidateSpan's own "unreadable" message says Invalid hours value '...' while the downstream McpHelpers.ValidateHoursBack message for out-of-range says Invalid hours_back value '...' — a small vocabulary seam between the two error paths. It's already called out in the doc comments as an accepted cosmetic inconsistency and is pinned explicitly by FleetSweepWebFeedTests, so this is a deliberate tradeoff, not an oversight — just flagging it exists in case that's not the intent.
  • BuildDetailAsync fetches verdicts and (conditionally) the would-have-paged ledger sequentially rather than concurrently. Given these are single-request, presentation-only reads against small per-sweep row counts, this is not a meaningful performance concern.

Nothing else stood out as a bug, drift, or regression. Tests (FleetSweepStateRungTests, FleetSweepWebFeedTests, DarlingWebOidcTests, DarlingWebAssetsTests) look like they actually pin the behaviors described (span refusal table, wire-shape honesty rules, seat matrix, frontend parity via source-scan) rather than just exercising happy paths.

🤖 Generated with Claude Code

@erikdarlingdata
erikdarlingdata merged commit fc0b399 into dev Sep 16, 2026
8 checks passed
@erikdarlingdata
erikdarlingdata deleted the feat/3466-sweep-web-feed branch September 16, 2026 03:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant