Skip to content

QuickJS engine: threshold-based VM-memory snapshotting (WORKFLOW_SNAPSHOT_THRESHOLD) - #3251

Merged
TooTallNate merged 6 commits into
mainfrom
quickjs-vm-threshold-snapshots
Oct 1, 2026
Merged

TooTallNate merged 6 commits into
mainfrom
quickjs-vm-threshold-snapshots

Conversation

@TooTallNate

@TooTallNate TooTallNate commented Jul 31, 2026 •

Copy link
Copy Markdown
Member

Note

Supersedes #3053. Stacked on #3250 (quickjs-vm-snapshots). Review only the commits after #3250's.

Summary

Experimental threshold-based VM-memory snapshotting for the QuickJS engine. Snapshots are taken only once WORKFLOW_SNAPSHOT_THRESHOLD events have been processed since the last one, not at every suspension:

  • Short-lived runs never snapshot. They keep pure replay with no snapshot round-trips.
  • Long/forever runs stop scaling their resume cost with event-log length. A resumption restores the VM heap and replays only the events recorded after the snapshot's cursor.

How it works

  • Policy: the WORKFLOW_SNAPSHOT_THRESHOLD env var (default 0, disabled) or per-run executionContext.snapshotThreshold, stamped at start() like WORKFLOW_VM. An invalid handler-side value disables snapshotting with a warning.
  • Save (suspension exit, threshold met): capture the VM memory. Then, after the response goes out (waitUntil):
    1. Frame the heap with its restore-relevant metadata, bound to the run id.
    2. Compress (zstd on the threadpool, gzip fallback) and encrypt with the run's key.
    3. world.experimental_snapshots.save.
      Captures over 32 MB plaintext are skipped, and the run is latched so later suspensions skip the capture.
  • Restore (later invocation):
    1. load, check format/engine version and bounds on the metadata.
    2. Decrypt, and decompress with a size cap.
    3. Verify the sealed metadata matches the envelope.
      The saved position is the read cursor plus the exact number of events it covers, so a restore adds only the events listed after it.
    4. Restore over the cached WASM module, re-register every host callback from one shared list, and reinstall process.env from the current invocation.
    5. Replay the delta.
      The load is skipped when no snapshot can exist yet (log below the threshold).
  • Encryption: runs without an encryption key are only snapshotted with WORKFLOW_SNAPSHOT_ALLOW_UNENCRYPTED=1. A run with a key only accepts snapshots encrypted with it.
  • Delete: after the run's terminal event, off the response path, whenever a snapshot may exist. A save that lands after the run finished deletes itself.
  • Fallback is always full replay. A missing, rejected, or unrestorable snapshot logs a warning and boots fresh against the full log. The log remains the source of truth.

Determinism model (restore + partial replay)

A resumption may restore a snapshot older than the log head and must re-derive everything in between deterministically:

  • Seeded PRNG: it stays on the run's base seed. The snapshot records how many draws the heap consumed (rngDraws), and restore fast-forwards that many. Ids are therefore position-based: identical across snapshot generations, identical to a no-snapshot full replay, and identical across concurrent resumes from different snapshots, so the world's dedup still collapses them.
  • ULID factory and clock: the correlation-id ULID factory state (lastUlid) and the deterministic clock's high-water mark (clockMs) are persisted and continued.
  • Re-fed events: feeding already-consumed events is harmless. Settled resolvers are gone, re-scanned terminals for consumed ids are dropped, and hook deliveries are deduped by eventId.

Validation

  • Unit tests with the real VM:
    • restore/resume, and id parity with full replay;
    • partial replay across multiple suspensions;
    • host-callback re-registration;
    • current process.env after restore;
    • no step results retained in the heap.
  • Differential fuzz (quickjs-snapshot-fuzz.test.ts): random parallel, sequential and raced step schedules driven live, through random restores (including older snapshots, lagging re-feeds and split bursts), and as one full replay, all of which must agree. A short run is in the unit suite; a nightly workflow runs more seeds and steps.
  • Real-VM entrypoint tests against an in-memory World (quickjs-snapshot-generations.test.ts):
    • MaxEventsExceeded fires exactly at the limit across many save/restore generations;
    • each saved cursor covers exactly its saved event count;
    • thresholds 1/3/4/1000 write the same log and result as no snapshots and leave nothing behind;
    • an older save landing after a newer one converges.
  • Entrypoint tests with snapshot storage mocked:
    • load gating;
    • delete on completion, on restore failure, and past the threshold;
    • rejection of plaintext, mismatched-metadata, other-run, out-of-bounds, and oversized snapshots.
  • CI: a quickjs-snapshot matrix leg (nextjs-turbopack, threshold=1) across the local dev/prod/postgres e2e jobs. It fails if the server log shows no snapshot restores.

Notes

  • Snapshot bytes are tied to the quickjs-wasi build that produced them. A mismatch is a clean miss. On Vercel, runs stay on their deployment, so deploys don't invalidate in-flight snapshots.
  • Docs: WORKFLOW_SNAPSHOT_THRESHOLD and WORKFLOW_SNAPSHOT_ALLOW_UNENCRYPTED in v5 Runtime Tuning, including what a snapshot contains and how it is protected.

Copilot AI review requested due to automatic review settings July 31, 2026 03:22
@TooTallNate
TooTallNate requested review from a team and ijjk as code owners July 31, 2026 03:22
@changeset-bot

changeset-bot Bot commented Jul 31, 2026 •

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: facf224

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 21 packages
Name Type
@workflow/core Patch
@workflow/world Patch
workflow Patch
@workflow/builders Patch
@workflow/cli Patch
@workflow/next Patch
@workflow/nitro Patch
@workflow/vitest Patch
@workflow/web-shared Patch
@workflow/web Patch
@workflow/world-testing Patch
@workflow/errors Patch
@workflow/world-local Patch
@workflow/world-postgres Patch
@workflow/world-vercel Patch
@workflow/astro Patch
@workflow/nest Patch
@workflow/rollup Patch
@workflow/sveltekit Patch
@workflow/vite Patch
@workflow/nuxt Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@vercel

vercel Bot commented Jul 31, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
example-nextjs-workflow-turbopack Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
example-nextjs-workflow-webpack Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
example-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workbench-astro-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workbench-express-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workbench-fastify-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workbench-hono-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workbench-nestjs-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workbench-nitro-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workbench-nuxt-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workbench-python-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workbench-sveltekit-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workbench-tanstack-start-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workbench-vite-workflow Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workflow-docs Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workflow-swc-playground Building Building Preview, v0 Oct 1, 2026 6:27pm UTC
workflow-tarballs Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC
workflow-web Ready Ready Preview, v0 Oct 1, 2026 6:27pm UTC

@github-actions

github-actions Bot commented Jul 31, 2026 •

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

✅ All tests passed

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • abortDeterministicBranchWorkflow: if-check takes same path on first-run and replay (nest · vercel-prod / vercel / node / production)
  • health check (CLI) - workflow health command reports healthy endpoints (fastify · local-dev / local / node / stable)
  • hookMinRetentionWorkflow - terminal Hook cannot resume and its token stays unavailable (example · vercel-ws-transport / vercel / node / ws)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • run-pickup-stall · runs app-generated source against a registered step (tanstack-start) · at 18:30:31Z · abandoned wrun_01M3WBPG1HAPYCHYM4P0F62G7P
  • cold-start-warmup · suite warmup (tanstack-start) · at 18:30:38Z · abandoned wrun_01M3WBPG1QPXT1PNV18ZCVFK08
  • run-pickup-stall · abortFromStepWorkflow: step abort cancels an in-flight sibling step (nuxt) · at 18:32:53Z · abandoned wrun_01M3WBTV6ZJEE8DME5B5ZP0AV5
  • run-pickup-stall · hookCleanupTestWorkflow - hook token reuse after workflow completion (nextjs-webpack) · at 18:37:10Z · abandoned wrun_01M3WC2NVGYA26S0M9BP32PNMS
  • run-pickup-stall · hookGetConflictWorkflow - awaiting hook.getConflict() registers hook without payload (nextjs-webpack) · at 18:37:19Z · abandoned wrun_01M3WC2YV0NRPT973CJ4AJ64ZF
  • run-pickup-stall · 'hookGetConflictWithPriorStepWorkflow' - hook.getConflict() does not block step execution (nextjs-webpack) · at 18:37:20Z · abandoned wrun_01M3WC2ZGJ7JC9P9GAB29SKXP7

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3904 0 875 4779
✅ 💻 Local Development 4582 0 551 5133
✅ 📦 Local Production 4582 0 551 5133
✅ 🐘 Local Postgres 4582 0 551 5133
✅ 🪟 Windows 342 0 12 354
✅ 🌐 Cross-language Conformance 68 0 84 152
✅ dynamic-runs 0 0 0 0
✅ vercel-http-transport 879 0 183 1062
✅ vercel-multi-region 27 0 0 27
✅ vercel-ws-transport 595 0 113 708
Total 19561 0 2920 22481
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 142 0 35
✅ astro-quickjs 142 0 35
✅ example-node 142 0 35
✅ example-quickjs 142 0 35
✅ express-node 142 0 35
✅ express-quickjs 142 0 35
✅ fastify-node 142 0 35
✅ fastify-quickjs 142 0 35
✅ hono-node 142 0 35
✅ hono-quickjs 142 0 35
✅ nest-node 142 0 35
✅ nest-quickjs 142 0 35
✅ nextjs-turbopack-node 169 0 8
✅ nextjs-turbopack-quickjs 169 0 8
✅ nextjs-webpack-node 169 0 8
✅ nextjs-webpack-quickjs 169 0 8
✅ nitro-node 142 0 35
✅ nitro-quickjs 142 0 35
✅ nuxt-node 142 0 35
✅ nuxt-quickjs 142 0 35
✅ python-node 66 0 111
✅ sveltekit-node 161 0 16
✅ sveltekit-quickjs 161 0 16
✅ tanstack-start-node 142 0 35
✅ tanstack-start-quickjs 142 0 35
✅ vite-node 142 0 35
✅ vite-quickjs 142 0 35

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 148 0 29
✅ astro-stable-quickjs 148 0 29
✅ express-stable-node 148 0 29
✅ express-stable-quickjs 148 0 29
✅ fastify-stable-node 148 0 29
✅ fastify-stable-quickjs 148 0 29
✅ hono-stable-node 148 0 29
✅ hono-stable-quickjs 148 0 29
✅ nest-stable-node 148 0 29
✅ nest-stable-quickjs 148 0 29
✅ nextjs-turbopack-canary-node 176 0 1
✅ nextjs-turbopack-canary-quickjs 176 0 1
✅ nextjs-turbopack-quickjs-snapshot 176 0 1
✅ nextjs-turbopack-stable-node 176 0 1
✅ nextjs-turbopack-stable-quickjs 176 0 1
✅ nextjs-webpack-canary-node 176 0 1
✅ nextjs-webpack-canary-quickjs 176 0 1
✅ nextjs-webpack-stable-node 176 0 1
✅ nextjs-webpack-stable-quickjs 176 0 1
✅ nitro-stable-node 148 0 29
✅ nitro-stable-quickjs 148 0 29
✅ nuxt-stable-node 148 0 29
✅ nuxt-stable-quickjs 148 0 29
✅ sveltekit-stable-node 167 0 10
✅ sveltekit-stable-quickjs 167 0 10
✅ tanstack-start-node 148 0 29
✅ tanstack-start-quickjs 148 0 29
✅ vite-stable-node 148 0 29
✅ vite-stable-quickjs 148 0 29

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 148 0 29
✅ astro-stable-quickjs 148 0 29
✅ express-stable-node 148 0 29
✅ express-stable-quickjs 148 0 29
✅ fastify-stable-node 148 0 29
✅ fastify-stable-quickjs 148 0 29
✅ hono-stable-node 148 0 29
✅ hono-stable-quickjs 148 0 29
✅ nest-stable-node 148 0 29
✅ nest-stable-quickjs 148 0 29
✅ nextjs-turbopack-canary-node 176 0 1
✅ nextjs-turbopack-canary-quickjs 176 0 1
✅ nextjs-turbopack-quickjs-snapshot 176 0 1
✅ nextjs-turbopack-stable-node 176 0 1
✅ nextjs-turbopack-stable-quickjs 176 0 1
✅ nextjs-webpack-canary-node 176 0 1
✅ nextjs-webpack-canary-quickjs 176 0 1
✅ nextjs-webpack-stable-node 176 0 1
✅ nextjs-webpack-stable-quickjs 176 0 1
✅ nitro-stable-node 148 0 29
✅ nitro-stable-quickjs 148 0 29
✅ nuxt-stable-node 148 0 29
✅ nuxt-stable-quickjs 148 0 29
✅ sveltekit-stable-node 167 0 10
✅ sveltekit-stable-quickjs 167 0 10
✅ tanstack-start-node 148 0 29
✅ tanstack-start-quickjs 148 0 29
✅ vite-stable-node 148 0 29
✅ vite-stable-quickjs 148 0 29

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 148 0 29
✅ astro-stable-quickjs 148 0 29
✅ express-stable-node 148 0 29
✅ express-stable-quickjs 148 0 29
✅ fastify-stable-node 148 0 29
✅ fastify-stable-quickjs 148 0 29
✅ hono-stable-node 148 0 29
✅ hono-stable-quickjs 148 0 29
✅ nest-stable-node 148 0 29
✅ nest-stable-quickjs 148 0 29
✅ nextjs-turbopack-canary-node 176 0 1
✅ nextjs-turbopack-canary-quickjs 176 0 1
✅ nextjs-turbopack-quickjs-snapshot 176 0 1
✅ nextjs-turbopack-stable-node 176 0 1
✅ nextjs-turbopack-stable-quickjs 176 0 1
✅ nextjs-webpack-canary-node 176 0 1
✅ nextjs-webpack-canary-quickjs 176 0 1
✅ nextjs-webpack-stable-node 176 0 1
✅ nextjs-webpack-stable-quickjs 176 0 1
✅ nitro-stable-node 148 0 29
✅ nitro-stable-quickjs 148 0 29
✅ nuxt-stable-node 148 0 29
✅ nuxt-stable-quickjs 148 0 29
✅ sveltekit-stable-node 167 0 10
✅ sveltekit-stable-quickjs 167 0 10
✅ tanstack-start-node 148 0 29
✅ tanstack-start-quickjs 148 0 29
✅ vite-stable-node 148 0 29
✅ vite-stable-quickjs 148 0 29

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 171 0 6
✅ nextjs-turbopack-quickjs 171 0 6

✅ 🌐 Cross-language Conformance

App Passed Failed Skipped
✅ python 68 0 84

✅ dynamic-runs

App Passed Failed Skipped
✅ astro-vercel 0 0 0
✅ example-vercel 0 0 0
✅ express-vercel 0 0 0
✅ fastify-vercel 0 0 0
✅ hono-vercel 0 0 0
✅ nest-vercel 0 0 0
✅ nextjs-turbopack-vercel 0 0 0
✅ nextjs-webpack-vercel 0 0 0
✅ nitro-vercel 0 0 0
✅ nuxt-vercel 0 0 0
✅ sveltekit-vercel 0 0 0
✅ tanstack-start-vercel 0 0 0
✅ vite-vercel 0 0 0

✅ vercel-http-transport

App Passed Failed Skipped
✅ example 142 0 35
✅ express 142 0 35
✅ hono 142 0 35
✅ nextjs-turbopack 169 0 8
✅ nitro 142 0 35
✅ vite 142 0 35

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

✅ vercel-ws-transport

App Passed Failed Skipped
✅ example 142 0 35
✅ express 142 0 35
✅ nextjs-turbopack 169 0 8
✅ vite 142 0 35

📋 View full workflow run

@VaguelySerious VaguelySerious left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI review: blocking issues found

Comment thread packages/core/src/runtime/quickjs-entrypoint.ts
Comment thread packages/core/src/runtime/quickjs-runtime.ts Outdated
Comment thread packages/core/src/runtime/quickjs-runtime.ts Outdated
Comment thread packages/core/src/runtime/quickjs-entrypoint.ts Outdated
Comment thread packages/core/src/runtime/quickjs-entrypoint.ts Outdated
Comment thread packages/core/src/runtime/quickjs-entrypoint.ts Outdated
Comment thread packages/core/src/runtime/quickjs-entrypoint.ts Outdated

@pranaygp pranaygp left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed the incremental diff (12 files, +926) and ran the stack top locally. Not merge-ready — headline finding inline at the load gate: snapshots never restore on world-postgres or world-vercel, and nothing in CI can tell, because a snapshot that never restores falls back to full replay, which is correct behavior. The quickjs-snapshot CI legs are currently red only on the inherited #3049 overflow bug; once that's fixed they'd go green while snapshots remain pure cost on both production worlds.

Local validation of what does work (world-local): snapshot lifecycle is real — 101 snapshot_saved / 283 snapshot_load diag checkpoints across a full e2e leg, 4 MB-ish .bin/.json pairs created and deleted on completion, restored: true resumptions observed. The position-based PRNG fast-forward design is correct: the restored-run-produces-identical-correlationIds property is the right one to pin, and I could not construct a divergence; the per-(runId, correlationId) dedup backstop can't wedge. Host-callback re-registration via the unified list is complete (no other newFunction sites). The eventsCursor frontier is sound across all three worlds (strictly exclusive cursors, no off-by-one). Compression is already workerd-safe.

One n=1 observation from a full snapshot-leg run worth your eyes: parallelStepsThenWebhookWorkflow - no hook_conflict from same-tick replay race failed once with a correctness assertion (harness posted body-<tokenA> for a token it listed from storage, workflow expected body-<tokenB>), i.e. a post-restore disagreement about the run's own webhook token. 10/10 passes in isolation afterwards and absent from a second full leg — but a restore-path token/PRNG-position determinism bug is what it would smell like if it recurs.

Non-blocking: first-invocation captures with no cursor are taken then dropped; the restore drain loop lacks the boot path's iteration-bound warning; deleteSnapshotIfAny early-returns when the threshold is 0, so flipping the env off strands stored snapshots (and a waitUntil save landing after a terminal delete re-creates one); __hookPayloadBuffer.__processedEventIds grows monotonically in-heap (snapshot size for long-lived reusable hooks only increases); __wdk_env isn't refreshed after restore; the docs say invalid threshold values "throw at startup" but getSnapshotThresholdFromEnv throws on first invocation. Changeset omits @workflow/world (this PR modifies packages/world/src/snapshots.ts).

Stack coordination: #3263 rebases under this per its own description — do that first. Three silent semantic breaks to check on the rebase: the inline __generateUlid registration won't be re-registered on restore (the branch's own comment says host callbacks must never be inline); the restore path's "serde survives in the heap" comment becomes false (serde is host-side after #3263 — createQuickJSSerde must run on restore); and the ULID monotonic factory's state moves host-side, so it's no longer captured by the snapshot and the rngDraws fast-forward desynchronizes.

const version = loaded.metadata.formatVersion;
if (
(version !== undefined && version !== SNAPSHOT_FORMAT_VERSION) ||
loaded.metadata.rngDraws === undefined

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: this gate rejects every snapshot on world-postgres and world-vercel, so the feature never restores on either production world. Neither world persists the new metadata fields: world-postgres writes/reads only eventsCursor + createdAt (migration 0018 has no columns for eventCount/rngDraws/formatVersion), and world-vercel sends/parses only the two original headers (carrying the new fields also needs a workflow-server change). Only world-local round-trips them, which is why the local tests pass.

Net effect on prod worlds: every qualifying suspension pays session.snapshot() (two full heap copies) + compression + encryption + a 5–15 MB PUT, and every resume throws it away and full-replays — strictly worse than WORKFLOW_SNAPSHOT_THRESHOLD=0. Persisting the fields in both worlds is the right fix; tolerating missing rngDraws is not safe (a restored heap with rngDraws: 0 re-draws and collides correlationIds).

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in eee5695 (+ 374dbaa on the base PR). Rather than adding postgres columns and a workflow-server change, the base PR now packs the FULL metadata object into the stored blob itself (snapshot envelope: metadata + bytes in one self-describing unit, schema-validated on decode with passthrough for forward compat). postgres stores the envelope in the existing data column; vercel sends it as the PUT/GET body the server already stores opaquely — so every metadata field this PR adds (eventCount, rngDraws, formatVersion, and the newer serdeRootPtr/lastUlid/clockMs/engineVersion) round-trips on BOTH production worlds today, no server change needed, and future fields need no storage work at all. The gate still refuses to tolerate missing rngDraws (agreed that direction is unsafe); it just no longer has a reason to fire on prod worlds. Also: world.snapshots went optional on the base PR, so the entrypoint now feature-detects and forces the threshold to 0 when absent.

// its next qualifying suspension instead.)
const snapshot = capturedSnapshot;
const totalEventCount = restoredEventCount + seenEventIds.size;
safeWaitUntil(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"Off the response path" is false on the default codec: this IIFE runs synchronously up to its first real await, and compress with the default write codec is zlib.zstdCompressSync — synchronous compression of a multi-MB heap image blocks the event loop before the response flushes. The gzip fallback (CompressionStream) is genuinely async, so this only bites the default path. Defer past the current tick (Promise.resolve().then(...)) or use the async codec for snapshots.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in eee5695, both halves: (1) the waitUntil IIFE now yields past the current tick (setImmediate) before touching the bytes, so the response flushes first; (2) compress() gained a preferAsync option — zstd via node:zlib's callback API (libuv threadpool, identical output bytes; sync fallback where unavailable) — and the snapshot pipeline uses it, so the compression itself no longer blocks the loop regardless of when it starts. Event-payload compression keeps the sync path (small payloads, unchanged behavior).

// (pre-snapshot count persisted in the metadata + delta) — otherwise a
// run that keeps snapshotting would never accumulate enough delta to
// trip the ceiling it exists to enforce.
const restoredEventCount = existingSnapshot?.metadata.eventCount ?? 0;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

restoredEventCount double-counts after a restore failure. The restore-failure catch resets existingSnapshot/lastEventsCursor to fall back to full replay but can't reset this const; the ceiling check then adds the stale count to a seenEventIds set that now holds the entire log, so MaxEventsExceededError can fire well below the real limit — and the next save stamps the inflated eventCount, compounding. Currently masked by the load-gate finding (no restores on prod worlds ⇒ no failures to fall back from), which is exactly the latent-bug shape that surfaces the day that's fixed.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in eee5695: restoredEventCount is a let and the restore-failure fallback resets it to 0 alongside existingSnapshot/lastEventsCursor — from that point events/seenEventIds cover the whole run, so the ceiling compares the true total and the next save stamps an accurate eventCount.

) {
try {
capturedSnapshot = session.snapshot();
if (capturedSnapshot.data.byteLength > MAX_SNAPSHOT_PLAINTEXT_BYTES) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ceiling enforced after the expensive work: session.snapshot() has already made two full copies of the WASM heap by this check, and since WASM linear memory never shrinks, a run that once crossed 32 MB re-pays both copies at every subsequent suspension and discards the result every time. Gate on VM memory size before snapshotting, or latch a per-run "too big" flag on first rejection.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in eee5695 with the latch you suggested: a process-local bounded set of runIds whose heap exceeded the ceiling. Since WASM linear memory never shrinks, once-oversized is always-oversized — later suspensions of that run now skip BEFORE session.snapshot() (the two heap copies), not after. Process-local is the right scope: the warm instance replaying the same run repeatedly is where the repeated cost lived; a cold instance pays one probe and re-latches.

Comment thread packages/world/src/snapshots.ts Outdated
* Current snapshot format version, bumped when the heap layout or the
* metadata contract changes incompatibly.
*/
export const SNAPSHOT_FORMAT_VERSION = 1;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

formatVersion covers the SDK's metadata envelope, not the heap image — the QJSS header that deserializeSnapshot validates is identical across quickjs-wasi builds, so bytes from one build restored by another pass validation and execute as undefined behavior in the interpreter. Real deploy-skew hazard (a quickjs-wasi bump mid-rollout has live snapshots from the old build), and a data-corruption-class failure. Cheap close: the library exposes vm.versions — persist it in metadata and reject a mismatch. Related smaller gap: the deterministic clock's high-water mark (vmNowMs) isn't persisted either, so Date.now() in a restored VM can regress until the first event advances it.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in eee5695, both parts. Engine pin: the asset build script now bakes the installed quickjs-wasi version into quickjs-assets.generated.ts (the exact build the embedded WASM came from — more precise than runtime vm.versions probing and available before any VM exists); it's stamped into metadata as engineVersion and the load gate treats any mismatch/absence as a clean miss, so a mid-rollout quickjs-wasi bump degrades to full replay instead of restoring a foreign heap. Clock: clockMs (the deterministic clock's high-water mark at capture) persists in metadata and primes the restored VM via the monotonic advanceClock, so Date.now() no longer regresses to run-creation time until the first delta event.

// (pre-snapshot count persisted in the metadata + delta) — otherwise a
// run that keeps snapshotting would never accumulate enough delta to
// trip the ceiling it exists to enforce.
const restoredEventCount = existingSnapshot?.metadata.eventCount ?? 0;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

restoredEventCount is captured as a const before snapshot restore and cannot be reset when restore fails and the code falls back to full event-log replay, causing the pre-snapshot count to be added on top of the now-full events/seenEventIds.

Fix on Vercel

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI: Fixed earlier in eee5695 (restoredEventCount is a let and gets reset to 0 on the restore-failure fallback). The fix carries into the rebased history as 083fe1c.

// its next qualifying suspension instead.)
const snapshot = capturedSnapshot;
const totalEventCount = restoredEventCount + seenEventIds.size;
safeWaitUntil(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Snapshot compression claimed to run off the response path via safeWaitUntil actually runs synchronously on the response-path call stack for the default zstd codec, blocking the event loop while compressing a multi-MB heap.

Fix on Vercel

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI: Fixed earlier in eee5695. The waitUntil pipeline yields past the current tick (setImmediate) before touching the bytes, and compress runs with preferAsync, which uses zstd on the libuv threadpool. The fix carries into the rebased history as 083fe1c.

@github-actions

github-actions Bot commented Aug 14, 2026 •

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 42 total

fence=per-spec

scenario outcome events virt replay violations
✅ smoke-no-steps completed 3 0ms ok 0
✅ smoke-one-step completed 6 0ms ok 0
✅ hook-at-step-started completed 12 0ms ok 0
✅ hook-at-step-completed completed 12 0ms ok 0
✅ hook-at-hook-created completed 12 0ms ok 0
✅ deadline-hook-wins completed 7 1.0h ok 0
✅ deadline-expires completed 7 1.0h ok 0
✅ step-vs-timer-early-settlement completed 8 1.0h ok 0
✅ long-sleep completed 11 30.0d ok 0
✅ hook-never-arrives stalled 3 0ms skipped 0
✅ step-retries-twice completed 10 2.0s ok 0
✅ parallel-steps completed 9 0ms ok 0
✅ hook-on-execution-state completed 12 0ms ok 0
✅ peek-hook-before-branch completed 12 0ms ok 0
✅ peek-hook-after-branch completed 12 0ms ok 0
✅ peek-hook-at-registration completed 12 0ms ok 0
✅ race-hook-before-probe completed 12 0ms ok 0
✅ race-hook-after-probe completed 12 0ms ok 0
✅ race-duplicate-delivery completed 13 0ms ok 0
✅ attr-hook-before-step completed 11 0ms ok 0
✅ attr-hook-after-step completed 11 0ms ok 0
✅ attr-from-step-body completed 13 0ms ok 0
✅ fork-hook-after-timeout completed 14 1.0m ok 0
✅ fork-hook-before-timeout completed 14 1.0m ok 0
✅ count-hook-after-timeout completed 17 1.0m ok 0
✅ count-hook-before-timeout completed 20 1.0m ok 0
✅ stale-read-step-count-fork completed 20 1.0m ok 0
✅ stale-read-equal-step-counts completed 14 1.0m ok 0
✅ step-vs-step-fork completed 12 0ms ok 0
✅ step-vs-step-fork-fenced completed 12 0ms ok 0
✅ fence-catches-benign-direction completed 12 5ms ok 0
✅ in-flight-before-decision completed 17 1.0m ok 0
❌ in-flight-before-decision-counted completed 17 1.0m ok 0
✅ in-flight-after-decision completed 19 2.0m ok 0
✅ stale-read-step-count-fork-fenced completed 20 1.0m ok 0
✅ fork-hook-wins completed 13 1.0m ok 0
✅ fork-timeout-wins completed 13 1.0m ok 0
✅ unclaimed-payload-under-fork completed 17 1.0m ok 0
✅ claimed-payload-under-fork completed 17 1.0m ok 0
✅ writers-independent-step-bodies completed 12 0ms ok 0
✅ writers-scripted-tempo completed 12 0ms ok 0
✅ cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

@karthikscale3 karthikscale3 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review from a DynamoDB read-throttling investigation on workflow-server — I came at this PR asking whether it would relieve the event-log read amplification we're seeing in production. Short answer: yes, this targets exactly the right thing. Notes below are prefixed AI (found independently) or AI+Human (a colleague pointed me at the area).

Context for the numbers cited: event-log reads are ~89% of base-table throttle events, read throttling occurs in 168 of 168 hourly buckets over a 7-day window, and the worst run I traced had 6,198 events with ~175 concurrent resumes replaying from the identical cursor. This PR's "replay only the delta since the snapshot cursor" is the right lever for that.

} | null = null;
if (
snapshotsStorage &&
snapshotThreshold > 0 &&

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI+Human — possible regression for short runs

The guards here are "snapshots enabled" and "not the first invocation", but nothing checks whether this run has ever crossed the threshold. So a run with threshold=100 that only ever accumulates 20 events issues an awaited, blocking GET on every resume for a snapshot that cannot exist.

The miss isn't cheap either: server-side loadSnapshot goes straight to S3.GetObjectCommand and lets the 404 throw (I confirmed the 404 path in an integration test against the real endpoints). So each wasted probe is a network hop plus an S3 round-trip on the critical path of a resume.

This is the case the threshold design explicitly protects — "Short-lived runs below the threshold never pay the snapshot cost" — and the save path honors it while the load path doesn't.

Two options:

  • Gate on the loaded event count, which the runtime already has: skip the load when count < threshold, since no snapshot can exist by definition.
  • Or stamp snapshotSaved: true into executionContext on the first save.

Worth a unit test asserting snapshots.load is not called below the threshold.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI: Agreed, this was a real regression against the threshold design. Fixed in 6503a34. The load is now skipped whenever no snapshot can exist yet. A snapshot is only saved once the saving VM has processed at least threshold events, so a log below the threshold guarantees a miss:

  • Complete preload below the threshold: a caller-attested complete preloaded log (the lazy hook fast path) shorter than the threshold skips the probe outright.
  • Process-local latch: when an invocation suspends with no snapshot and a total below the threshold, the run goes into a bounded, globalThis-held set, and the next invocation in that process skips the probe. The only way this goes stale is if another instance pushed the log past the threshold and saved a snapshot in the meantime. Skipping the load then just costs one full replay, which is always correct, and that invocation sees a log at or above the threshold and clears the entry.

I didn't stamp this into executionContext: it's fixed at start(), and a runtime write to it would add a round-trip of its own. Cold instances still pay one probe per run, but warm instances replaying a short run don't. Tests in quickjs-snapshot-load-gate.test.ts check that snapshots.load isn't called below the threshold and is still called at or above it.

version !== SNAPSHOT_FORMAT_VERSION ||
loaded.metadata.rngDraws === undefined ||
loaded.metadata.serdeRootPtr === undefined ||
loaded.metadata.engineVersion !== quickjsWasiVersion

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI — correlated invalidation on a quickjs-wasi bump

The pin itself is right, and the rationale in the comment is convincing (identical QJSS heap-image headers across builds ⇒ restoring under the wrong build is undefined behavior). No argument with rejecting here.

The concern is that this failure is correlated across runs: a deploy that bumps quickjs-wasi invalidates every in-flight long run's snapshot simultaneously, and they all revert to full-log replay at once.

That spike lands precisely on the long runs this feature exists to protect, and it lands on a table where read throttling already occurs in 168/168 hourly buckets. A synchronized fallback-to-full-replay is close to the worst possible shape for it.

Not asking for a fix in this PR, but it needs a deploy story — staggered invalidation, a grace window that accepts the previous engine version, or at minimum a documented deploy note.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI: Agreed this needs a deploy story. Where it bites:

  • On Vercel: runs keep executing on the deployment they started on, so bumping quickjs-wasi in a new deploy doesn't touch in-flight runs' snapshots. Old runs keep restoring on the old build, and new runs start on the new one. No synchronized invalidation.
  • Where it does happen: environments that serve in-flight runs from new code (self-hosted rolling updates on Postgres, say). There, every in-flight run falls back to full replay at its next resume, once, then re-snapshots at its next qualifying suspension.

I've documented this in the WORKFLOW_SNAPSHOT_THRESHOLD section of Runtime Tuning (6503a34). A grace window that accepts the previous engine version isn't safe: the QJSS heap-image header is identical across builds, so a cross-build restore is undefined behavior, not a slower path. Staggering invalidation would need something like jittered snapshot-age expiry, which is doable as a follow-up if non-Vercel deployments turn out to need it.

Copy link
Copy Markdown
Member Author

AI: I rebased this onto the synced #3250, which is on the latest main. Merge commits can't be signed from this environment, and the repository requires signed commits. The branch is now two commits on top of #3250:

  • 083fe1c: the PR's content squashed and synced. One conflict resolution changes behavior: main added a __validateAttributeWrite host callback inline in the fresh-boot path, which the snapshot-restore path would not have re-registered. It now lives in the shared hostCallbacks list, and there's a regression test: restore, then an oversized attribute write gets the size error. The snapshot eventsCursor is still tracked separately from main's QuickJSLogView read cursor, and it only advances on listed pages that are fully fed to the VM.
  • 6503a34: review follow-ups. snapshots.load is skipped when no snapshot can exist yet, the snapshot latches are moved onto globalThis to satisfy main's module-scope-state rule, and there's a docs note on engine-bump invalidation.

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor
Framework Flow route Step reg. Framework output
hono 267.2 KiB (+1.2 KiB) 97.3 KiB (+53 B) 2.00 MiB (+10.1 KiB)
nextjs-turbopack 274.6 KiB (+1.2 KiB) 426 B (±0) 985.3 KiB (+1.8 KiB)
About these numbers

Sizes are gzip; parentheses show the change against main.
Flow route and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.

facf224 · run

@pranaygp pranaygp left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed this together with #3250 (storage API) and #3253 (dry run). Blast radius for users who don't opt into QuickJS/snapshots is minimal (only the start() env parse/stamp, the additive Postgres migration, and additive exports), but the feature itself has security and correctness gaps worth fixing before merge. Each bug below was reproduced with a scratch test against this branch.

Correctness bugs

1. Snapshots keep every step result forever, so snapshotting switches itself off for long runs. (quickjs-runtime.ts:2193-2211 together with :2361-2383; the restore loop at :1769 has the same shape)

continueWithEvents loops over the same batch while it makes progress. On the second pass, events that were just consumed have no resolver, so processEvents stores their results in __terminalBuffer[cid]. Those correlation IDs are never used again, so the entries are never removed.

Without snapshots this memory dies with the invocation. With snapshots the heap lives on, gets saved, and the leak compounds. Measured with 20 KB step results:

after step 1 50 200 400
snapshot size 1.77 MB 2.75 MB 5.83 MB 9.96 MB

That's about 20 KB per step, and a full replay of the same log leaves the buffer empty. A run with 20 KB results hits the 32 MB cap after roughly 1,500 steps, and the oversized-snapshot latch then turns snapshotting off for that run — exactly the long-running runs this feature is for. It also means every old step result stays in stored snapshots. Fix: don't buffer a terminal event for a correlation ID that has already been consumed (keep a set of consumed IDs in the VM), and add a test that snapshot size stays flat over N steps.

2. After a restore fails, the snapshot is never deleted. (quickjs-entrypoint.ts:1522, :2442)

The fallback sets existingSnapshot = null. If the run then completes or fails in that same invocation, deleteSnapshotIfAny exits early and the bad snapshot stays in storage for good. Reproduced: delete called 0 times.

3. A snapshot can be left behind when the run finishes. (:2369-2425 vs :2442)

Invocation A suspends and schedules its save via waitUntil, which runs after the response goes out. Invocation B picks the run up right away (inline steps make this the normal pattern), finds no snapshot yet, completes the run, and skips the delete because it neither restored nor saved one. A's save then lands on a finished run. Same thing when the per-process "below threshold" cache is stale.

Fix: delete whenever restoredEventCount + seenEventIds.size >= snapshotThreshold (a snapshot can only exist past the threshold), and have save refuse terminal runs or rely on a storage-side TTL/cascade. Local and Postgres have no TTL at all.

4. A restored run sees process.env from when its snapshot chain started. (quickjs-runtime.ts:1919)

installProcessEnv only runs on a fresh boot, so process.env travels inside the heap. After a restore, workflow code sees the environment from the original boot — rotated secrets and env-based config changes never show up. Full replay and the Node engine read the current env on every invocation, so snapshotting changes what the workflow observes. Fix: reinstall process.env on restore, or back it with a host callback.

5. PR description is out of date. It says the PRNG seed mixes in eventsCursor; the code now fast-forwards by rngDraws (the right fix). Please update it.

Security

Before this PR, write access to world storage could only forge data. A snapshot is executable state (bytecode, closures, pending continuations), so write access to snapshot storage now means running arbitrary code in the workflow VM — code that can call host callbacks and request any registered step with any arguments, and steps run with full host privileges.

S1. When the run has a key, a plaintext snapshot is still accepted. (quickjs-entrypoint.ts:1233)

decrypt() returns anything without an encr prefix unchanged, even when a key is configured. Reproduced: with a key set, a plaintext load() result reached startQuickJSWorkflow as the snapshot to restore. For events that's tolerable; for a heap it bypasses encryption to run code. When encryptionKey is set, reject snapshots that aren't encrypted.

S2. Snapshot metadata isn't authenticated.

The metadata sits in plaintext inside the storage envelope and nothing ties it to the encrypted bytes:

  • Metadata from one snapshot can be paired with the bytes of another from the same run; the restore then uses the wrong cursor and silently replays from the wrong log position (what #3250's envelope was built to prevent).
  • rngDraws has no maximum (world/src/snapshots.ts:31) and the fast-forward loop at quickjs-runtime.ts:1572 runs synchronously: 20M draws blocked for 1.4 s in my test; Number.MAX_SAFE_INTEGER hangs the invocation until timeout, on every retry.
  • serdeRootPtr is an arbitrary pointer handed to adoptSerdeRoot.

Fix: pass runId, canonical metadata and engineVersion as AES-GCM associated data, and cap rngDraws and similar fields.

S3. No decompression size limit on load. (:1237)

decompress() inflates with no output cap, so a zstd bomb inside a 64 MB envelope can OOM the function instance. The 32 MB cap only applies on save; enforce it on load too (e.g. zstd maxOutputLength).

S4. Secrets and step data stored unencrypted on local and Postgres.

Only world-vercel implements getEncryptionKeyForRun. On world-local (.workflow-data/snapshots/*.snapshot) and world-postgres (workflow_snapshots.data), the snapshot is plaintext apart from compression. It contains the handler's whole process.env (database URLs, API keys, …) plus leaked step results (bug 1), and because of bugs 2 and 3 it can outlive the run with no TTL. Suggestions: don't store env in the heap (fixing bug 4 fixes this), refuse to snapshot without an encryption key unless explicitly opted in, and say this plainly in runtime-tuning.mdx.

S5. Rename the storage API to experimental_snapshots. (world/src/interfaces.ts:513)

snapshots? on Storage plus public exports (SnapshotMetadataSchema, encode/decodeSnapshotEnvelope, SNAPSHOT_FORMAT_VERSION) signal to community worlds that this is a stable contract to implement. Please rename to experimental_snapshots and mark the exports @experimental (or keep them internal). Also settle all persisted names now, since renaming later needs a data migration: the executionContext.snapshotThreshold key saved on runs, the Postgres table workflow_snapshots, and maybe WORKFLOW_EXPERIMENTAL_SNAPSHOT_THRESHOLD.

Risks and missing tests

  • The save/restore/delete lifecycle is only tested with the VM mocked. Nothing covers save metadata, the restore-failure fallback, format/engine mismatch, delete on completion/failure/run-gone, or MaxEventsExceeded with a restored count. Please add real-VM tests on world-local for bugs 1–3 and S1–S3.
  • Older snapshots can overwrite newer ones. Overlapping waitUntil saves are last-write-wins — still correct, but resume cost regresses. Consider "only replace if eventCount is higher".
  • Memory per save. Each save holds several full-heap copies at once (plaintext, compressed, encrypted, envelope), up to 32 MB plaintext per run; many runs saving concurrently on one instance adds up.
  • A bad handler env value breaks every QuickJS invocation. An invalid WORKFLOW_SNAPSHOT_THRESHOLD on the handler makes getSnapshotThreshold (vm-mode.ts) throw every time → retry loop. Warn and treat as 0 on the handler side instead.
  • The run-completed write waits on the snapshot delete (:2461, :2742). Consider deleting after the terminal event, or via waitUntil.
  • #3253 is stale: it carries migration 0019_add_snapshots_table.sql while #3250 has 0024. Rebase before trusting its CI signal.

Observability

  • Snapshot activity only shows up in debug-level diagnostics and a few warn logs; there are no span attributes. Suggest something like workflow.quickjs.snapshot.{restored, delta_events, restore_ms, fallback_reason} on the invocation span.
  • Put a separate span around the waitUntil save pipeline (save_ms, plaintext_bytes, stored_bytes); it runs after the response and is invisible today.
  • A 100% miss rate should be visible as a metric, not just warn logs.
  • world-postgres snapshots storage isn't wrapped in instrumentObject, unlike local and vercel.

Docs

  • runtime-tuning.mdx: mark WORKFLOW_SNAPSHOT_THRESHOLD experimental; document that snapshots contain process.env and step data and are unencrypted on worlds without an encryption key; document restored-env behaviour (bug 4) and the missing TTL on local/Postgres. Line 223 still calls snapshotting "future".
  • Changesets: mention the new Postgres table in the world-postgres changeset, and use the renamed field.

Nits

  • getSnapshotThresholdFromEnv accepts "1e3", "0x10" and " 5 "; use /^\d+$/.
  • quickjs-runtime.ts: the JSDoc for restoreWorkflowVM now sits above incrementUlidRandom, leaving two doc blocks stacked.
  • The restore-fallback full-log fetch duplicates the earlier pagination loop; pull it into a helper.
  • :1184 uses the inline type import('@workflow/world').SnapshotMetadata; import SnapshotMetadata as a named type like quickjs-runtime.ts does.

Recommendation: request changes — safe for existing users, but please fix S1–S3 and bugs 1–4 and land the experimental_snapshots rename first.

Copy link
Copy Markdown
Member Author

AI: Thanks for the thorough pass. Both branches are rebased on the latest main, and the fixes are in d87ef61 (this PR) and 423a253 (#3250). Every bug and security item below has a regression test, and each one fails without its fix.

Correctness

  1. Retained step results. Fixed. The VM now records every correlation id it constructed an awaiting promise for (__awaitedCids). A re-scanned terminal that finds no resolver for such an id has already been consumed, so it is dropped instead of buffered. The same fix covers the restore loop, since it goes through processEvents. A new real-VM test runs 60 steps with 20 KB results and asserts __terminalBuffer is empty after each one; without the fix it held 59 entries. I also tried a snapshot byte-size assertion, but the raw heap image size stayed the same with or without the fix in my runs, so the test asserts on the buffer directly.
  2. Snapshot not deleted after a failed restore. Fixed. Terminal cleanup keys on "a snapshot may exist": one was found (usable or not), this invocation saved one, or the log has reached the threshold. The restore-failure fallback no longer clears that.
  3. Snapshot left behind when the run finishes. Fixed along the lines you suggested:
    • Terminal paths delete whenever restoredEventCount + seenEventIds.size >= threshold. The delete runs after the terminal event is written (including the already-finished conflict branches), via waitUntil, so run_completed no longer waits on it.
    • After a save lands, the saving invocation reads the run's status and deletes its own snapshot if the run already finished. Terminal paths write the terminal event before they delete, so any interleaving ends with the snapshot gone. The one case left is a run cancelled with no invocation observing it; world-local and world-postgres have no TTL for that, which the docs now say.
  4. Stale process.env after restore. Fixed. The restore path reinstalls process.env from the current invocation. Real-VM test: a restored run sees the value set after the snapshot was taken.
  5. PR description. Rewritten. It now describes the rngDraws fast-forward and the rest of the current design.

Security

  • S1 (plaintext accepted when the run has a key). Fixed. A run with an encryption key rejects any stored snapshot that isn't encr-encrypted. That includes encp sealed payloads, since anyone with the run's public key can produce those.
  • S2 (unauthenticated metadata). Fixed without needing AAD support in the encryption layer. The saved bytes are a frame: the restore-relevant metadata (runId, cursor, eventCount, rngDraws, lastUlid, serdeRootPtr, clockMs, engineVersion, formatVersion) in canonical form, followed by the heap, then compressed and encrypted as one payload. On load, the sealed metadata must equal the envelope's, or the snapshot is a miss. With a key this is authenticated: cross-snapshot pairing, other-run bytes, and edited envelope fields are all rejected. Without a key it still catches torn or mismatched writes.
    • rngDraws is capped at 10M and serdeRootPtr must be a 32-bit address. Both are checked on the envelope metadata before any decrypt or decompress work.
    • Format version bumped to 3, so older snapshots are a clean miss.
  • S3 (unbounded decompression). Fixed. decompress() takes a maxOutputBytes option: zstd uses maxOutputLength, gzip counts output as it streams. Snapshot loads cap at the 32 MB plaintext ceiling plus frame overhead. The test uses a correctly sealed snapshot of an over-ceiling heap that is under 1 MB on the wire.
  • S4 (plaintext snapshots on local/Postgres). Taken your "refuse without a key unless opted in" suggestion. Runs without an encryption key are never snapshotted unless the handler sets WORKFLOW_SNAPSHOT_ALLOW_UNENCRYPTED=1; the local/postgres CI snapshot legs set it. Reinstalling process.env alone isn't enough: the per-invocation env copy is still written into the heap, and freed WASM memory isn't zeroed, so encryption is the real protection. The docs now say what a snapshot contains and that storage write access means code execution.
  • S5 (rename). Done on Add experimental world.experimental_snapshots storage interface (local, postgres, vercel) #3250: Storage.experimental_snapshots, with @experimental on the interface and on every snapshot export from @workflow/world. I kept the persisted names as they are: executionContext.snapshotThreshold, the workflow_snapshots table, and WORKFLOW_SNAPSHOT_THRESHOLD. The docs mark the feature experimental. Happy to rename the env var or context key before merge if you'd rather they carry the prefix too.

Risks and observability

  • Lifecycle tests. New entrypoint tests cover:
    • restore plus delete on completion;
    • delete after a failed restore;
    • delete past the threshold with no snapshot found, and no delete below it;
    • the plaintext, metadata-mismatch, other-run, out-of-bounds and oversized rejections;
    • the unencrypted opt-in.
      Snapshot storage and the VM are mocked there. The real-VM coverage is in quickjs-runtime.test.ts (heap retention, process.env, host callbacks). I haven't added an end-to-end world-local test; the threshold=1 CI leg exercises that path.
  • Invalid handler env. getSnapshotThresholdForHandler warns once and treats an invalid value as 0. start() still rejects it up front. The parser only accepts /^\d+$/.
  • Observability.
    • The invocation span gets workflow.quickjs.snapshot.{restored, delta_events, restore_ms, fallback_reason}.
    • The save runs inside a workflow.quickjs.snapshot.save span with save_ms, plaintext_bytes and stored_bytes.
    • fallback_reason makes the miss rate queryable from traces. I haven't added a separate counter.
  • Not changed, deliberately:
    • Older-overwrites-newer needs a compare-and-set in each World's save. It affects cost only, not correctness, so I'd rather do it as a follow-up.
    • Peak memory per save is still several heap-sized copies, bounded by the 32 MB ceiling and the oversize latch.
    • world-postgres doesn't wrap any of its storage in instrumentObject today, so doing it for snapshots alone felt inconsistent. Worth a separate change for the whole World.
    • [DO NOT MERGE] CI dry-run: QuickJS as the default workflow VM engine #3253 is outside these two PRs.

Docs and nits

  • Docs. runtime-tuning.mdx marks the feature experimental and documents contents, encryption, the unencrypted opt-in, the restored-env behavior, and retention on local/Postgres. The stale "future" line is fixed.
  • Changesets. This PR's changeset now includes @workflow/world and the renamed field, and Add experimental world.experimental_snapshots storage interface (local, postgres, vercel) #3250's names the new Postgres table and migration.
  • Nits. All four are done:
    • the /^\d+$/ threshold parse;
    • the restoreWorkflowVM JSDoc is back above its function;
    • the full-log fetch is one listRunLogFrom helper, and the restore fallback now also updates the log view;
    • SnapshotMetadata is imported as a named type.

Copy link
Copy Markdown
Contributor

Follow-up on testing, re-checked against the current head (d87ef61). Thanks for the fast turnaround: the rename, sealed metadata, plaintext rejection, inflate/rngDraws caps, env reinstall, heap-retention fix and lifecycle deletes all look right, and src/runtime is green here (62 files / 1086 tests).

The VM layer holds up well. What's thin is CI coverage of long runs, bursts at the threshold, the event limit and races, plus one remaining bug.

Remaining bug: the saved event position lags, so the event limit trips early

lastEventsCursor only advances on listed pages (quickjs-entrypoint.ts:1814). Events fed from write responses (takeQueuedEvents) are counted in seenEventIds and therefore in the saved eventCount, but they don't move the cursor. On restore, everything after the lagging cursor is fetched again and counted a second time (restoredEventCount + events.length, :1484/:1854). The error compounds with each save/restore generation, so a long run on Vercel (the only world with a maxEvents ceiling) gets MaxEventsExceeded before it reaches the real limit.

In a local e2e run at threshold=3 (on the previous head; this code path is unchanged), 12 of 18 stored snapshots had a cursor behind their eventCount, e.g. count 13 with the cursor at event 6. Fix: advance the persisted cursor for write-response events too, or derive eventCount from the cursor position rather than from seenEventIds.

What CI covers today

  • One snapshot lane (nextjs-turbopack, threshold=1, local dev/prod/postgres) running the existing e2e suite. The PR adds no e2e workflows.
  • Nothing asserts snapshots are actually used: a 100% miss rate would still pass. Locally only 119 of 361 resumes restored.
  • No Vercel lane sets the threshold, so the encrypted path and the maxEvents ceiling are never exercised end to end.
  • The e2e runs are short (a few to ~120 events). No long runs, many-generation runs, threshold bursts or event-limit tests.
  • CI is red (Vercel Prod cells, Node and QuickJS, on the experimental_force hook test, plus Windows unit tests; Add experimental world.experimental_snapshots storage interface (local, postgres, vercel) #3250 has the same). Probably not this PR, but it needs triage.

Tests to add

  1. Commit the differential fuzz test below (a quick version in unit tests, and a nightly run with more seeds and steps). It drives the same random schedule of parallel, sequential and race steps three ways: a live VM, a VM restored at random points (sometimes from an older snapshot, with a lagging cursor that re-feeds consumed events, and bursts split across the restore), and one clean full replay of the final log. It asserts identical step correlation IDs, step inputs and results. Here it passes 60 seeds × 30 steps and 4 seeds × 300 steps (1,784 restores). It does catch real breakage: FUZZ_MUTATE=rng (off-by-one rngDraws) and FUZZ_MUTATE=ulid (dropped lastUlid) both fail it.

  2. Entrypoint tests with a real VM and an in-memory world that can inject cursor lag and slow saves. Cover bursts of threshold−1 / threshold / threshold+1 events, two invocations overlapping a waitUntil save, an older save landing after a newer one, and MaxEventsExceeded firing exactly at the limit across many save/restore generations (this is the test that catches the cursor bug above).

  3. A long-run e2e workflow: ~2,000 sequential steps, a burst of 1,000 hook payloads, cancel/fail variants. Run at thresholds 1, 1000 and off, and assert the same result and event sequence as the Node engine, at least one restore, no snapshots left after the run finishes, no premature MaxEventsExceeded, and a time budget.

  4. Make the snapshot lane fail when it saw zero restores or left snapshots behind (count restored: true in the server log and files in workflow-data/snapshots).

  5. A Vercel lane with snapshots on: the only encrypted path, and the only one with an event limit.

  6. A snapshot variant of the replay benchmark to choose the default threshold. Measured here (single sequential-step workflow):

    events restore (decompress + restore) full replay
    21 20 ms 11 ms
    201 16 ms 31 ms
    2,001 24 ms 565 ms
    6,001 29 ms 4,154 ms

    Restore wins from roughly 100–200 events before network cost. The real app heap is ~15 MB (3.7 MB compressed), so very low thresholds cost more than they save on short runs.

On going past the event limit

Today the limit is still enforced (correctly, apart from the cursor bug). If we later want snapshotted runs to exceed it, every failure path currently falls back to a full replay, which is exactly what becomes infeasible at that size: an engine bump, a corrupt/missing snapshot, a load error, or the 32 MB size cap would leave a very long run stuck. That would need an exact cursor, a fallback other than full replay (e.g. keep the last K snapshots and fail loudly), and tests at 10–100× the current limit.

packages/core/src/runtime/quickjs-snapshot-fuzz.test.ts (differential fuzz, ready to commit)

Knobs: FUZZ_TRIALS (default 40), FUZZ_STEPS (default 30), FUZZ_MUTATE=rng|ulid (mutation check; drop that line before committing if you prefer).

/**
 * Review scratch: differential fuzz of snapshot/restore vs. live vs. full
 * replay. Same event schedule; snapshot driver restores at random points,
 * from random OLDER snapshots, with random cursor lag (re-feeding
 * already-consumed events) and random burst sizes.
 */
import seedrandom from 'seedrandom';
import { describe, expect, it } from 'vitest';
import { deserialize, serialize } from '../serialization/workflow-vm.js';
import { startQuickJSWorkflow } from './quickjs-runtime.js';

const run = {
  runId: 'wrun_01JXT21Q00AAAAAAAAAAAAAAAA',
  deploymentId: 'dpl_test',
  workflowName: 'w',
  input: undefined,
  status: 'running' as const,
  output: undefined,
  error: undefined,
  completedAt: undefined,
  startedAt: new Date('2025-01-01T00:00:00Z'),
  createdAt: new Date('2025-01-01T00:00:00Z'),
  updatedAt: new Date('2025-01-01T00:00:00Z'),
  specVersion: 2,
};
const N = Number(process.env.FUZZ_STEPS ?? 30);
const code = `
  var s = globalThis[Symbol.for("WORKFLOW_USE_STEP")]("step//test//s");
  async function workflow() {
    var acc = [];
    for (var i = 0; i < ${N}; i++) {
      if (i % 5 === 0) {
        var xs = await Promise.all([s(i, 0), s(i, 1), s(i, 2)]);
        acc.push(xs.join('+'));
      } else if (i % 7 === 0) {
        var w = await Promise.race([s(i, 'a'), s(i, 'b')]);
        acc.push('race:' + w);
      } else {
        acc.push(await s(i, Math.random()));
      }
    }
    return { acc: acc, t: Date.now(), r: Math.random() };
  }
  workflow.workflowId = "workflow//test//workflow";
  globalThis.__private_workflows.set("workflow//test//workflow", workflow);
`;
const opts = { workflowCode: code, workflowId: 'workflow//test//workflow', workflowRun: run };
const hex = (u?: Uint8Array) => (u ? Buffer.from(u).toString('hex') : '');

type Ev = any;
type Snap = { data: Uint8Array; meta: any; frontier: number };

async function drive(seed: string, mode: 'live' | 'snap') {
  const sched = seedrandom(seed); // schedule decisions (identical across modes if cids agree)
  const chaos = seedrandom(seed + ':chaos'); // snapshot-driver-only decisions
  const log: Ev[] = [
    { eventId: 'e_rc', runId: run.runId, eventType: 'run_created', eventData: { input: serialize([]) }, createdAt: new Date(Date.UTC(2025, 0, 1)) },
  ];
  const trace: string[] = [];
  const created = new Set<string>();
  let t = 1;
  let frontier = 0; // log index the live VM has consumed up to
  const snaps: Snap[] = [];
  let restores = 0;
  let session = await startQuickJSWorkflow({ ...opts, events: log.slice() as any });
  frontier = log.length;
  let result = session.result;
  for (let round = 0; round < 10_000 && result.suspended; round++) {
    const fresh = result.suspended.pendingOperations.filter(
      (o: any) => o.type === 'step' && !created.has(o.correlationId)
    ) as any[];
    for (const o of fresh) trace.push(`${o.correlationId}|${hex(o.input)}`);
    // Complete a random-size burst of pending (not yet completed) steps.
    const pending = result.suspended.pendingOperations.filter(
      (o: any) => o.type === 'step' && !log.some((e) => e.eventType === 'step_completed' && e.correlationId === o.correlationId)
    ) as any[];
    if (pending.length === 0) throw new Error('stuck: suspended with nothing to complete');
    const k = 1 + Math.floor(sched() * pending.length);
    // Shuffle completion order deterministically.
    const order = pending.slice().sort(() => sched() - 0.5).slice(0, k);
    for (const o of pending) {
      if (created.has(o.correlationId)) continue;
      created.add(o.correlationId);
      log.push({ eventId: `c_${o.correlationId}`, runId: run.runId, eventType: 'step_created', correlationId: o.correlationId, eventData: { stepName: 'step//test//s' }, createdAt: new Date(Date.UTC(2025, 0, 1, 0, 0, t++)) });
    }
    for (const o of order) {
      log.push({ eventId: `d_${o.correlationId}`, runId: run.runId, eventType: 'step_completed', correlationId: o.correlationId, eventData: { result: `R(${o.correlationId.slice(-4)})` }, createdAt: new Date(Date.UTC(2025, 0, 1, 0, 0, t++)) });
    }

    if (mode === 'snap' && chaos() < 0.5) {
      // Snapshot the live VM (at its current frontier), then drop it and
      // restore from a random saved snapshot (possibly OLDER), feeding the
      // delta with a random lag that re-feeds consumed events.
      const c = session.snapshot();
      snaps.push({ data: c.data, frontier, meta: { eventsCursor: 'x', createdAt: new Date(), rngDraws: c.rngDraws + (process.env.FUZZ_MUTATE === 'rng' ? 1 : 0), lastUlid: process.env.FUZZ_MUTATE === 'ulid' ? undefined : c.lastUlid, serdeRootPtr: c.serdeRootPtr, clockMs: c.clockMs, engineVersion: c.engineVersion } });
      session.dispose();
      const pick = snaps[Math.floor(chaos() * snaps.length)];
      const lag = Math.floor(chaos() * 6);
      const from = Math.max(1, pick.frontier - lag);
      // Burst split: restore with part of the delta, continue with the rest.
      const delta = log.slice(from);
      const cut = Math.floor(chaos() * (delta.length + 1));
      session = await startQuickJSWorkflow({ ...opts, events: delta.slice(0, cut) as any, existingSnapshot: { data: pick.data, metadata: pick.meta } });
      restores++;
      result = session.result;
      if (result.suspended && cut < delta.length) result = await session.continueWithEvents(delta.slice(cut) as any);
    } else {
      result = await session.continueWithEvents(log.slice(frontier) as any);
    }
    frontier = log.length;
  }
  if (!result.completed) throw new Error('did not complete: ' + JSON.stringify(Object.keys(result)));
  const out = deserialize(result.completed.result);
  session.dispose();
  return { out, trace, log, restores };
}

describe('snapshot differential fuzz', () => {
  const TRIALS = Number(process.env.FUZZ_TRIALS ?? 40);
  it(`${TRIALS} seeds: snapshot driver == live driver == full replay`, async () => {
    let totalRestores = 0;
    for (let i = 0; i < TRIALS; i++) {
      const seed = `seed-${i}`;
      const live = await drive(seed, 'live');
      const snap = await drive(seed, 'snap');
      totalRestores += snap.restores;
      expect(snap.trace, seed).toEqual(live.trace);
      expect(snap.out, seed).toEqual(live.out);
      // Oracle: a single fresh full replay of the final log.
      const full = await startQuickJSWorkflow({ ...opts, events: snap.log as any });
      expect(full.result.completed, seed).toBeDefined();
      expect(deserialize(full.result.completed!.result), seed).toEqual(snap.out);
      full.dispose();
    }
    process.stderr.write(`FUZZ trials=${TRIALS} steps=${N} restores=${totalRestores}\n`);
  }, 600_000);
});

Copy link
Copy Markdown
Member Author

AI: Thanks for re-checking. The cursor bug is fixed, most of the test gaps are filled, and both branches are rebased on the latest main (this PR now ends at ecaf7fa; #3250 is at a839548).

The cursor bug. Confirmed and fixed. I reproduced it first: in a real-VM entrypoint test with maxEventsLimit: 40, the limit tripped once the log reached 28 events.

  • QuickJSLogView now tracks how many events its read cursor covers. It advances on listing pages (counting every listed event, seen or not) and on complete inline deltas (by the delta's length), and becomes unknown when events arrive with no cursor after them.
  • A snapshot saves that pair: eventsCursor is the view's cursor and eventCount is exactly the events it covers. A restore adds only the events listed after it, so events fed from write responses are counted once, and the lagging-cursor double count can't happen.
  • Saves are skipped while the position isn't exact, or while events are still queued for the VM, so the cursor never covers events the heap doesn't have.
  • The restore-failure fallback resets the position to the full re-read.
  • The test now passes. It checks that the limit fires only once the log has reached it, and that every saved cursor's slot equals its saved eventCount.

Tests added

  1. Differential fuzz: committed as quickjs-snapshot-fuzz.test.ts, tidied and typed.
    • The unit suite runs 20 seeds × 30 steps (about 5 s); FUZZ_TRIALS and FUZZ_STEPS override that.
    • A new nightly QuickJS Snapshot Fuzz workflow (also manually triggerable) runs 200 × 300 by default.
    • I kept FUZZ_MUTATE. rng and ulid both fail it here, as you found.
    • It also asserts that the snapshot driver really restored.
  2. Real-VM entrypoint tests (quickjs-snapshot-generations.test.ts): an in-memory World with slot ids, cursored listings and inline deltas on writes, so events do reach the VM from write responses. It covers:
    • the exact event limit across ~15 save/restore generations, plus cursor/count agreement on every save;
    • thresholds 1, 3, 4 and 1000 against a no-snapshot baseline: identical log shape (event types and correlation ids) and decrypted result, restores actually happening, nothing saved at 1000, and no snapshot left after completion. Thresholds 3 and 4 give bursts that straddle the threshold;
    • a save held until after the next invocation's save lands, so an older snapshot overwrites a newer one; the run still converges with the baseline.
  3. Snapshot lane must restore: restores now log at info, and the snapshot legs set DEBUG=workflow:runtime:info. A new step after the dev/prod/postgres e2e runs counts restored VM snapshot lines in the server log and fails on zero. I didn't add a "no snapshots left" check: the suite has runs that legitimately never finish (parked hooks, externally cancelled runs), and those keep their snapshot by design, as the docs note.
  4. Threshold guidance: your numbers are now in the docs as guidance on choosing a value: restore starts to pay off at a few hundred events, and very low values mostly add saves. I haven't added a snapshot variant to the benchmark suite yet.

Not done in this pass: a ~2,000-step / 1,000-hook-payload e2e workflow (3), and a Vercel lane with snapshots on (5).

  • They're the right next step, but both add new workbench workflows and CI lanes across every framework, so I'd rather land them as a follow-up.
  • The Vercel lane in particular should go in together with a restore check that reads spans or logs from the deployment, which the local log grep can't do.
  • Until then, the real-VM tests above run the encrypted path end to end: they use a run key, so restores go through the authenticated, encrypted snapshot codec.

CI triage. On the previous heads:

On going past the event limit: agreed on all counts. This PR keeps the limit enforced, now exactly. Exceeding it would need a fallback other than full replay, retained snapshot generations, and tests at 10–100× the limit, and belongs in its own design.

Comment thread .changeset/quickjs-threshold-snapshots.md Outdated
vercel Bot and others added 6 commits October 1, 2026 11:24
…SHOT_THRESHOLD)

Squashed content of #3251, synced with main and #3250. Conflict
resolution with main's QuickJS changes:
- Hoist main's __validateAttributeWrite host callback into the shared
  hostCallbacks list so the snapshot-restore path re-registers it, and pass
  historicalAttributeIds to the restored live session.
- Keep the snapshot eventsCursor tracking alongside main's QuickJSLogView
  read cursor.

Co-Authored-By: Nathan Rajlich <71256+TooTallNate@users.noreply.github.com>
…ld, document engine-bump invalidation

- Skip world.snapshots.load when no snapshot can exist yet: a complete
  preloaded log shorter than the threshold, or a run this process last saw
  suspend below it (process-local, staleness only costs a full replay).
- Hold the snapshot latches on globalThis via globalSingleton (module-scope
  state rule from main).
- Docs: note that an engine-changing upgrade serving in-flight runs from new
  code makes their snapshots fall back to full replay at once.

Co-Authored-By: Nathan Rajlich <71256+TooTallNate@users.noreply.github.com>
- Stop buffering re-scanned terminals whose resolver already settled, so a
  long-lived (snapshotted) heap no longer retains every step result.
- Seal the restore-relevant metadata, bound to the run id, inside the
  encrypted snapshot and require it to match the envelope on load; reject
  plaintext snapshots for runs with an encryption key; cap decompression
  and rngDraws/serdeRootPtr before doing restore work. Format version 3.
- Only snapshot runs without an encryption key when
  WORKFLOW_SNAPSHOT_ALLOW_UNENCRYPTED is set.
- Reinstall process.env on restore so restored runs see the current env.
- Delete snapshots after the terminal event whenever one may exist (found,
  saved, restore failed, or log past the threshold), off the response path;
  a save that lands after the run finished removes itself.
- An invalid handler-side threshold disables snapshotting with a warning.
- Snapshot span attributes and a span around the post-response save.
- Docs: mark experimental, describe snapshot contents, encryption and
  retention; changeset covers @workflow/world.

Co-Authored-By: Nathan Rajlich <71256+TooTallNate@users.noreply.github.com>
…re checks in CI

- Save the read cursor together with the number of events it covers
  (tracked by QuickJSLogView, including inline deltas), instead of a
  listing-only cursor paired with the count of every event seen. A lagging
  cursor made each restore count the lagging span twice, compounding per
  generation and tripping MaxEventsExceeded early. Saves are skipped while
  the position isn't exact or events are still queued for the VM.
- Real-VM entrypoint tests against an in-memory World: the event limit
  fires exactly at the limit across many generations, thresholds 1/3/4/1000
  write the same log and result as no snapshots and leave nothing behind,
  and an older save landing after a newer one converges.
- Differential fuzz of snapshot/restore vs live vs full replay (short run
  in unit tests, nightly workflow with more seeds and steps).
- Snapshot CI legs log restores and fail when there were none.
- Docs: guidance on choosing a threshold.

Co-Authored-By: Nathan Rajlich <71256+TooTallNate@users.noreply.github.com>
Co-Authored-By: Nathan Rajlich <71256+TooTallNate@users.noreply.github.com>
Signed-off-by: Peter Wielander <mittgfu@gmail.com>
@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

No backport to stable for b79ad01 (AI decision).

This commit adds an entirely new experimental capability — threshold-based VM-memory snapshotting for the QuickJS engine — gated by the new WORKFLOW_SNAPSHOT_THRESHOLD and WORKFLOW_SNAPSHOT_ALLOW_UNENCRYPTED env vars, a new snapshot codec, new SnapshotMetadata fields and format version, new telemetry attributes, and a new CI matrix leg. That is feature work, not a stability fix, so it belongs only on main.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

b79ad013d254d60cc1e53e65dfd806da24a8964b

@github-actions github-actions Bot mentioned this pull request Oct 1, 2026

This branch had an error being deployed

1 failed (outdated) and 17 active (16 outdated) deployments
Preview – workflow-docs — facf2249 Deployed Oct 1, 2026 by vercel[bot]
Preview – workflow-swc-playground — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workbench-nuxt-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – example-nextjs-workflow-webpack — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – example-nextjs-workflow-turbopack — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workbench-nestjs-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – example-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workbench-sveltekit-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workbench-tanstack-start-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workbench-vite-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workbench-astro-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workbench-nitro-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workbench-hono-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workbench-fastify-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workbench-express-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workflow-tarballs — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workflow-web — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Preview – workbench-python-workflow — 3cb39094 Deployed Aug 14, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants