Skip to content

Use worker snapshots for relay storage metrics - #7845

Merged
ravarora2 merged 1 commit into
mainfrom
codex/storage-snapshot-default
Sep 23, 2026
Merged

ravarora2 merged 1 commit into
mainfrom
codex/storage-snapshot-default

Conversation

@ravarora2

@ravarora2 ravarora2 commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Make the relay read completed storage-accounting snapshots from PostgreSQL by default. An environment with no snapshot emits no storage metrics. When a worker publishes its first result, the relay picks it up on the next successful leader usage tick, without a restart or configuration change.

This changes how storage usage is measured and reported. S3 remains the object store. PostgreSQL stores the completed counts, byte totals, community breakdown, and calculation metadata—not the objects themselves.

Why move relay storage accounting entirely to database reads?

Today the relay supports two sources: its own S3 scan (inline, the default) and the worker's saved snapshot (external). That requires operators to deploy a worker and then coordinate a separate relay-mode change in each environment.

The worker already owns the S3 listing and calculation. It runs in a separate process with its own object cap, memory limit, and deadline. Keeping a second scan inside the serving relay retains the resource pressure and failure modes that this isolation is meant to remove.

PostgreSQL provides a durable handoff. The worker replaces one completed snapshot atomically after a successful calculation. The relay reads that row and exposes its values through the existing metrics endpoint. A relay restart or leadership change can load the last completed result without scanning the bucket again.

There is deliberately no S3 fallback when a snapshot is absent or a read fails. Such a fallback would put the expensive scan back in the relay precisely when a worker is absent or unhealthy. It would also leave two calculation paths with different resource limits and freshness behavior.

What happens if no worker exists?

A worker that has never published is a normal, successful inactive state. The snapshot table must exist through the normal database schema, but it may contain zero rows. The reader returns success, emits no storage totals or failed-load gauge, and checks again on the next usage tick. It does not exit the relay, start a worker, or list S3.

The relay checks for a completed database row; it does not inspect Kubernetes Jobs or CronJobs. Worker presence and snapshot presence are therefore different:

Situation Storage-reader behavior
No worker and no snapshot row Return success, emit no storage metrics, and query again next tick.
Worker starts later and publishes its first row Read and emit it on the next successful leader usage tick; no restart or mode change.
Worker is removed or fails, but its last valid row remains Continue reporting the last completed totals with their increasing age.
Query succeeds but a previously present row has been removed Clear the cache and stop refreshing storage series; existing exporter series expire through the idle timeout.
Database read fails, times out, or returns an invalid snapshot Log a warning, set buzz_storage_snapshot_load_ok=0, retain any last-good cached totals, and retry next tick.
Snapshot table is missing Treat this as a schema error, not normal worker absence; follow the failed-read behavior above.
BUZZ_STORAGE_METRICS=off Skip the snapshot query and all storage metric emission.

No row means “not measured,” not zero bytes. A valid completed snapshot of an empty bucket can report zero. If a read fails before any good snapshot has been cached, only failed-load health is emitted.

These reads happen in the existing background metrics task, not the relay startup path. A storage-read failure does not request process exit or change readiness directly. A wider database outage can still affect the relay's existing database-dependent operations and readiness.

Read, parse, and emit path

Worker: S3 listing -> BucketSnapshot -> JSON -> PostgreSQL
Relay: PostgreSQL -> JSON -> BucketSnapshot -> existing metric gauges
Collector: relay /metrics endpoint -> Datadog
  1. main.rs::run_storage_sweep_tick calls the reader from the existing leader-only usage task. The default interval remains 300 seconds.
  2. storage_sweep.rs::run_storage_metrics_tick coordinates the read, cache update, failure health, and metric emission.
  3. refresh_persisted_snapshot calls the existing Db::load_storage_accounting_snapshot() method. A five-second timeout covers pool acquisition and the query. It uses the relay's existing pool.
  4. The existing query in buzz-db/src/store/storage_accounting.rs reads snapshot, completed_at, duration_ms, max_objects, and code_sha from the singleton row.
  5. serde_json::from_value(stored.snapshot) decodes the JSON into the existing buzz_media::BucketSnapshot type. Rust infers that type from cache_persisted_snapshot's argument. The generated Deserialize implementation handles all fields, including the UUID-keyed per_community map.
  6. emit_cached_storage_metrics sets the existing physical/logical totals and community byte/object gauges. The existing Prometheus exporter exposes them for Datadog collection; this code does not send snapshot JSON to Datadog.

The worker's BucketSnapshot/CommunityStorage structures, JSON field names, SQL publication method, and database schema are unchanged. The old external-mode JSON conversion moves out of main.rs into the reader helper. No new format or hand-written field parser is needed. Duration and object-cap metadata are validated before replacing the cache.

The five-second relay read timeout is separate from the worker's 30-second startup acquisition budget introduced in #7770. This PR does not change the worker's pool or startup logic.

Configuration and code changes

Configuration Before After
Unset, inline, or on Relay-local S3 scan Read completed database snapshots. inline logs one migration warning.
external or snapshot Read completed database snapshots Continue reading snapshots.
off Disabled Disabled.
Unknown or empty value Disabled with a configuration error Same.

Existing inline settings intentionally become reader aliases. Merely changing the default would leave older deployments on their explicit inline setting and preserve the need for coordinated configuration PRs.

Remove relay-local scan scheduling, in-flight scan state, scan configuration, and obsolete scan-attempt tests and health metrics. Preserve the physical/logical total and per-community metric names, leader-only emission, community scope filtering, and cleanup of old community labels.

The shared relay-state field remains in place; its comment now describes a cached worker result. The Helm chart defaults to relayMode: external while leaving storageAccounting.enabled: false. Its rendering test verifies that the reader is configured even when no worker is created. The reader regression tests cover missing rows, later publication, updates, invalid data, timeout, removal, reactivation, and the explicit off switch. docs/storage-accounting.md describes this operating model.

Freshness and rollout impact

buzz_storage_snapshot_age_seconds uses the worker's original completion timestamp. Re-reading an old row never makes it fresh. buzz_storage_snapshot_load_ok=1 means the row was read and decoded; it does not prove that the worker's latest attempt succeeded.

Monitor snapshot age against the worker's schedule and allowed runtime, and monitor Job failures separately. The old relay scan-attempt health gauges are retired because the relay no longer runs those scans.

Deploy this relay version through the usual release process. Once it reaches an environment, workers can be enabled independently. Already-released charts that still set inline work with the new binary; future chart releases default explicitly to external. An intentional off override still requires an operator to enable the reader.

Environments with no completed snapshot lose relay-calculated storage totals after this upgrade. That is the intended tradeoff for removing scans from the serving process. Environments with a prior snapshot continue showing its last completed totals and age. This PR does not install workers in additional environments or deploy the new relay image.

Related issue

Related to #4601. Follows the worker startup fix in #7770. The separate fleet-usage collector work in #7176 overlaps this code and may need reconciliation when it merges.

Testing

Built and ran the relay locally against isolated PostgreSQL and Redis with the legacy inline setting, five-second usage ticks, and a 15-second metric idle timeout. The same process was exercised through these steps:

  1. Start with an empty snapshot table: relay becomes ready and exposes no storage series.
  2. Publish a snapshot containing 123 bytes: the community gauge becomes 123.
  3. Replace it with 456 bytes: the gauge becomes 456.
  4. Remove the row: storage series expire without fabricating a zero-byte measurement.
  5. Publish 789 bytes: metrics reactivate without restarting the relay; readiness remains healthy.

The local S3 endpoint recorded zero requests. The unrelated Git conformance startup probe was disabled for this isolated reader check; the test covers storage accounting, not every other use of S3 in the relay. No new relay image was deployed to staging or production during this test.

The broad local integration run also hit an existing p0_pool_acquisitions_use_typed_operation_pairs_without_other failure in buzz-db/tests/observability_source.rs. The same source-only test fails on unmodified main at 0ef7a2222; this PR changes neither the test nor its failing source.

The full local just ci run stopped at mobile native-asset setup: the objective_c build hook received no SDK path from xcrun. Mobile tests did not run.

Generated with Codex

Signed-off-by: Ravneet Arora <rarora@squareup.com>
@github-actions

Copy link
Copy Markdown

🔐 Codex Security Review

Status: review required for the current range.

The current range is 3b2e50b15c6afba3ff3b8fbe1bdf0a7a69d29f7b...f458f6b8ff90a355f242a8302994d144b3e44bef.
A new review must complete for this exact range. When manual authorization
is required, a Block organization member must comment exactly
@buzz-security-review f458f6b8ff90a355f242a8302994d144b3e44bef to authorize a new review.
Any previous review applies only to its recorded range.

@ravarora2
ravarora2 marked this pull request as ready for review September 23, 2026 16:25
@ravarora2
ravarora2 requested a review from a team as a code owner September 23, 2026 16:25
@ravarora2
ravarora2 merged commit 4142d2b into main Sep 23, 2026
91 checks passed
@ravarora2
ravarora2 deleted the codex/storage-snapshot-default branch September 23, 2026 18:39
wpfleger96 pushed a commit that referenced this pull request Sep 23, 2026
…rcement

* origin/main:
  fix: route databricks claude fqns to anthropic messages (#7829)
  feat(relay): add opt-in newest-first thread windows (#7823)
  refactor: move agent Git bootstrap into ACP harness (#7819)
  Use worker snapshots for relay storage metrics (#7845)

Signed-off-by: Hayt <211b96e6a2b7f45fd4047988976c7bbbeeda0c15f3ae7b32eec20834b5a55118@buzz.block.builderlab.xyz>
wpfleger96 pushed a commit that referenced this pull request Sep 23, 2026
…erative-labels

* origin/main:
  fix: route databricks claude fqns to anthropic messages (#7829)
  feat(relay): add opt-in newest-first thread windows (#7823)
  refactor: move agent Git bootstrap into ACP harness (#7819)
  Use worker snapshots for relay storage metrics (#7845)
  fix(hooks): strip repo-local git env from pre-push test lanes (#7841)

Signed-off-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
bradseiler pushed a commit that referenced this pull request Sep 23, 2026
…in-gate

* origin/main: (30 commits)
  feat(agents): humanize uncurated Databricks model ids with a label grammar (#7844)
  fix: route databricks claude fqns to anthropic messages (#7829)
  feat(relay): add opt-in newest-first thread windows (#7823)
  refactor: move agent Git bootstrap into ACP harness (#7819)
  Use worker snapshots for relay storage metrics (#7845)
  fix(hooks): strip repo-local git env from pre-push test lanes (#7841)
  fix(mobile-infra): render push grant lifetimes as decimal in chart 0.3.2 (#7820)
  fix(agent): preserve Databricks Opus UC reasoning and tool continuation (#7840)
  feat(canvas): add version history with atomic restore (#6780)
  chore(release): release Buzz Desktop version 0.5.24 (#7817)
  test(desktop): stabilize unread and audio release smoke fixtures (#7821)
  fix(desktop): remember Inbox unread-only choice (#7672)
  feat(mobile-infra): support development App Attest (#7744)
  docs(nip-fi): clarify federated identity amendments (#7803)
  fix(desktop): refresh channels after access-revoked closure (#7784)
  fix(desktop): bound startup request bursts and recover quota refusals (#7790)
  fix(audit): frame hash inputs with TLV (#7492)
  fix(admin): allow cold storage worker DB startup (#7770)
  feat(relay): add admin HTTP routes for member restriction management (#7302)
  fix(relay): fire kick live side effects at convergence; persist target; fence re-add race with held lock (#7298)
  ...

Signed-off-by: coder 1 <a93f3b1decd199cec83848f116ff60c20776cdd03861b9bba2610acae7c9eeeb@buzz.block.builderlab.xyz>
michaelneale added a commit that referenced this pull request Sep 24, 2026
* origin/main:
  feat(agents): humanize uncurated Databricks model ids with a label grammar (#7844)
  fix: route databricks claude fqns to anthropic messages (#7829)
  feat(relay): add opt-in newest-first thread windows (#7823)
  refactor: move agent Git bootstrap into ACP harness (#7819)
  Use worker snapshots for relay storage metrics (#7845)
  fix(hooks): strip repo-local git env from pre-push test lanes (#7841)
  fix(mobile-infra): render push grant lifetimes as decimal in chart 0.3.2 (#7820)
  fix(agent): preserve Databricks Opus UC reasoning and tool continuation (#7840)
  feat(canvas): add version history with atomic restore (#6780)
  chore(release): release Buzz Desktop version 0.5.24 (#7817)
  test(desktop): stabilize unread and audio release smoke fixtures (#7821)
  fix(desktop): remember Inbox unread-only choice (#7672)
  feat(mobile-infra): support development App Attest (#7744)
  docs(nip-fi): clarify federated identity amendments (#7803)
  fix(desktop): refresh channels after access-revoked closure (#7784)
  fix(desktop): bound startup request bursts and recover quota refusals (#7790)
  fix(audit): frame hash inputs with TLV (#7492)
  fix(admin): allow cold storage worker DB startup (#7770)

Signed-off-by: Michael Neale <michael.neale@gmail.com>
wpfleger96 pushed a commit that referenced this pull request Sep 24, 2026
…n-surface

* origin/main:
  fix(ci): consume the published MinIO image (#7870)
  fix(mobile): converge sidebar managers on relay head with resume re-read (#7806)
  fix(ci): bootstrap the reusable MinIO image in GHCR (#7869)
  Discover alternate Buzz ACP commands (#6948)
  fix(hooks): surface nextest failures and stale pnpm deps in pre-push (#7850)
  feat(acp): run one prepared task from a file or stdin (#7851)
  Fix mobile heart and warning emoji with native font fallback (#7842)
  chore(mesh): upgrade MeshLLM to 0.76.2 (#7559)
  feat(agents): humanize uncurated Databricks model ids with a label grammar (#7844)
  fix: route databricks claude fqns to anthropic messages (#7829)
  feat(relay): add opt-in newest-first thread windows (#7823)
  refactor: move agent Git bootstrap into ACP harness (#7819)
  Use worker snapshots for relay storage metrics (#7845)
  fix(hooks): strip repo-local git env from pre-push test lanes (#7841)
  fix(mobile-infra): render push grant lifetimes as decimal in chart 0.3.2 (#7820)

Signed-off-by: Duncan <dcfd242e557282d7a1e2cf2e6877522682f1e5c6156dc92ca7d90eaedd3b0f95@buzz.block.builderlab.xyz>
TheSentinel454 added a commit that referenced this pull request Sep 24, 2026
…undation-local

* origin/main: (50 commits)
  feat(desktop): relay admin console for the /api/admin/v1 operator surface (#4768)
  fix(mobile): keep retired sections manager out of successor cache (#7873)
  Select one feature flag provider at compile time (#7677)
  chore(release): release Buzz Desktop version 0.5.25 (#7867)
  fix(ci): consume the published MinIO image (#7870)
  fix(mobile): converge sidebar managers on relay head with resume re-read (#7806)
  fix(ci): bootstrap the reusable MinIO image in GHCR (#7869)
  Discover alternate Buzz ACP commands (#6948)
  fix(hooks): surface nextest failures and stale pnpm deps in pre-push (#7850)
  feat(acp): run one prepared task from a file or stdin (#7851)
  Fix mobile heart and warning emoji with native font fallback (#7842)
  chore(mesh): upgrade MeshLLM to 0.76.2 (#7559)
  feat(agents): humanize uncurated Databricks model ids with a label grammar (#7844)
  fix: route databricks claude fqns to anthropic messages (#7829)
  feat(relay): add opt-in newest-first thread windows (#7823)
  refactor: move agent Git bootstrap into ACP harness (#7819)
  Use worker snapshots for relay storage metrics (#7845)
  fix(hooks): strip repo-local git env from pre-push test lanes (#7841)
  fix(mobile-infra): render push grant lifetimes as decimal in chart 0.3.2 (#7820)
  fix(agent): preserve Databricks Opus UC reasoning and tool continuation (#7840)
  ...
Signed-off-by: tornquist <tornquist@squareup.com>
TheSentinel454 added a commit that referenced this pull request Sep 24, 2026
…t/osc-event-write-chokepoint

* commit 'a6a3032e446e6e66e8e41a229ef655ea79f36202': (50 commits)
  feat(desktop): relay admin console for the /api/admin/v1 operator surface (#4768)
  fix(mobile): keep retired sections manager out of successor cache (#7873)
  Select one feature flag provider at compile time (#7677)
  chore(release): release Buzz Desktop version 0.5.25 (#7867)
  fix(ci): consume the published MinIO image (#7870)
  fix(mobile): converge sidebar managers on relay head with resume re-read (#7806)
  fix(ci): bootstrap the reusable MinIO image in GHCR (#7869)
  Discover alternate Buzz ACP commands (#6948)
  fix(hooks): surface nextest failures and stale pnpm deps in pre-push (#7850)
  feat(acp): run one prepared task from a file or stdin (#7851)
  Fix mobile heart and warning emoji with native font fallback (#7842)
  chore(mesh): upgrade MeshLLM to 0.76.2 (#7559)
  feat(agents): humanize uncurated Databricks model ids with a label grammar (#7844)
  fix: route databricks claude fqns to anthropic messages (#7829)
  feat(relay): add opt-in newest-first thread windows (#7823)
  refactor: move agent Git bootstrap into ACP harness (#7819)
  Use worker snapshots for relay storage metrics (#7845)
  fix(hooks): strip repo-local git env from pre-push test lanes (#7841)
  fix(mobile-infra): render push grant lifetimes as decimal in chart 0.3.2 (#7820)
  fix(agent): preserve Databricks Opus UC reasoning and tool continuation (#7840)
  ...

Signed-off-by: tornquist <tornquist@squareup.com>

# Conflicts:
#	crates/buzz-db/src/store/event.rs
brow added a commit that referenced this pull request Sep 25, 2026
…ction

* origin/main: (21 commits)
  docs(vision): add /buzz/v1 read endpoints to the protocol contract (#7879)
  🤖 fix(justfile): point just staging at the current staging relay (#7881)
  fix(relay-admin): make thread deletions atomic and fence expired action leases under row lock (#7853)
  feat(desktop): relay admin console for the /api/admin/v1 operator surface (#4768)
  fix(mobile): keep retired sections manager out of successor cache (#7873)
  Select one feature flag provider at compile time (#7677)
  chore(release): release Buzz Desktop version 0.5.25 (#7867)
  fix(ci): consume the published MinIO image (#7870)
  fix(mobile): converge sidebar managers on relay head with resume re-read (#7806)
  fix(ci): bootstrap the reusable MinIO image in GHCR (#7869)
  Discover alternate Buzz ACP commands (#6948)
  fix(hooks): surface nextest failures and stale pnpm deps in pre-push (#7850)
  feat(acp): run one prepared task from a file or stdin (#7851)
  Fix mobile heart and warning emoji with native font fallback (#7842)
  chore(mesh): upgrade MeshLLM to 0.76.2 (#7559)
  feat(agents): humanize uncurated Databricks model ids with a label grammar (#7844)
  fix: route databricks claude fqns to anthropic messages (#7829)
  feat(relay): add opt-in newest-first thread windows (#7823)
  refactor: move agent Git bootstrap into ACP harness (#7819)
  Use worker snapshots for relay storage metrics (#7845)
  ...

Signed-off-by: Tom Brow <tomb@block.xyz>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants