Skip to content

[core] Fail dev HMR cleanup on a stranded step registration - #3682

Merged
VaguelySerious merged 3 commits into
mainfrom
peter/windows-step-registration-cleanup
Aug 25, 2026
Merged

[core] Fail dev HMR cleanup on a stranded step registration#3682
VaguelySerious merged 3 commits into
mainfrom
peter/windows-step-registration-cleanup

Conversation

@VaguelySerious

Copy link
Copy Markdown
Member

Problem

Three of the last 120 Tests runs had E2E Windows Tests (node) burn its full ~30-minute
window and get cancelled with every e2e test failing on
[world-local] Queue message failed (HTTP 500), where the handlerError is a Next.js
_error page rather than anything workflow-shaped:

job flow-route 500s stranded-import errors
96228253373 10,135 20,346
96214515530 10,218 10,886
96202855009 10,280 20,557

The flow route never answered a single request successfully in any of them. The dev server
log says why:

⨯ app/.well-known/workflow/v1/flow/__step_registrations.js:33:1
Module not found: Can't resolve '../../../../../workflows/source-map-warning-fixture.ts'

dev.test.ts > "should not log source map warnings for workflow node_modules imports"
writes workflows/source-map-warning-fixture.ts, which declares a 'use step' function,
and its afterEach deletes it. The generated __step_registrations.js imports every
discovered step file by path, so that deletion only stops breaking the flow route once a
rediscovery regenerates the file. When the watcher drops the unlink — which is what happens
on Windows — the generated file keeps importing a path that no longer exists, the flow route
stops compiling, and every workflow dispatch for the rest of the job gets a 500 pointing at
a fixture from a test that already passed.

The HMR log confirms nothing recovers it: 164 → 165 → 166 → 165 → 165 workflows, then
three skips and no further rediscovery. afterEach's prewarm() only waits for
workflow dev hmr: ready, printed once at startup, so nothing ever waited for the
post-deletion rediscovery in the first place.

The generated workflow bundle inlines workflow sources instead of importing them, so only
step-declaring fixtures can strand it. That is why new-workflow.ts leaks harmlessly in
the same runs while the source-map fixture is fatal.

The existing fail-fast guard cannot see this

The Windows job already probes the dev server before the long e2e run, precisely to avoid
burning the job window on a wedged server. It probes GET /api/chat, which compiles
independently of the flow route, so it reported -> 405 and let the suite run in all three
jobs.

Changes

  1. packages/core/e2e/dev.test.ts — after deleting fixtures, cleanup waits for the
    generated step registrations to stop importing them. If they still do when the budget
    runs out, the fixture contents are written back so the shared dev server keeps serving
    the rest of the suite, and the hook fails naming the stranded file. CI already skips the
    e2e suite when dev.test.ts fails, so this turns a 30-minute burn with ~10k misleading
    500s into an immediate, accurate failure. The convergence wait runs before the
    node_modules fixture-package removal so a restored fixture still resolves its import.

  2. .github/workflows/tests.yml — the pre-e2e probe now POSTs
    /.well-known/workflow/v1/flow?__health, the SDK's own health endpoint on the route
    that actually breaks. It returns 200 on a healthy dev server and would have returned 500
    in all three jobs above.

Validation

Run locally against workbench/nextjs-turbopack on macOS (all 9 dev tests, so including
the flow-route HMR fuzz cases that also delete step-declaring files):

  • Happy path: the convergence wait never trips on any of the tests that delete fixtures. A
    deletion converges in ~2s, so the watcher path itself is sound and the Windows miss is the
    discriminator.
  • should follow Next flow-route HMR rebuild rules for body-only changes fails on macOS
    (expected 2 to be 1 on an exact HMR log count). It fails identically with unmodified
    origin/main's dev.test.ts against the same dev server, so it is pre-existing and
    unrelated. That case is skipped on Windows.
  • Forced-failure path (convergence budget shortened so it always expires): the hook fails
    with the stranded path named, the fixture is restored, and
    POST /.well-known/workflow/v1/flow?__health still answers 200.
  • POST /.well-known/workflow/v1/flow?__health returns 200 on a healthy dev server and 500
    when a path reachable from __step_registrations.js does not resolve, while
    GET /api/chat keeps answering 405 — the asymmetry the second change fixes.

The underlying Windows watcher miss is worth chasing separately; this PR stops it from
costing a whole job and makes it name itself.

Note on a related hole, not changed here

The Linux e2e-local-dev lane runs pnpm vitest run packages/core/e2e/dev.test.ts; sleep 10
and then continues into the full e2e suite, so a dev.test.ts failure there is discarded.
That lane would swallow the new hook failure too. Tightening it would newly red the lane on
pre-existing dev.test.ts flakes (the flow-route HMR fuzz case already fails on macOS
against unmodified main, see below), so it seemed worth raising separately rather than
bundling in.

`dev.test.ts` fixtures that declare a `'use step'` function are imported by
path from the generated `__step_registrations.js`, so deleting one only stops
breaking the flow route once a rediscovery regenerates that file. When the dev
server's watcher drops the unlink, the generated file keeps importing a path
that no longer exists: the flow route stops compiling and every later workflow
dispatch in the job gets a 500 that names a fixture from a test that already
passed. Three of the last 120 `Tests` runs lost the whole ~30-minute Windows
e2e window this way, each with ~10k flow-route 500s and no successful dispatch.

Cleanup now waits for the generated step registrations to drop every file it
deleted. If they still reference one when the budget expires, the fixture is
written back so the shared dev server keeps serving the rest of the suite, and
the hook fails naming the stranded path. The wait runs before the node_modules
fixture-package removal so a restored fixture still resolves its import.

The pre-e2e probe moves from `GET /api/chat` to
`POST /.well-known/workflow/v1/flow?__health`. /api/chat compiles independently
of the flow route, so it answered 405 in all three jobs while every dispatch
was getting a 500.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@changeset-bot

changeset-bot Bot commented Aug 19, 2026

Copy link
Copy Markdown

🦋 Changeset detected

Latest commit: d11c0dd

The changes in this PR will be included in the next version bump.

This PR includes changesets to release 16 packages
Name Type
@workflow/core Patch
@workflow/builders Patch
@workflow/cli Patch
@workflow/next Patch
@workflow/nitro Patch
@workflow/vitest Patch
@workflow/web-shared Patch
@workflow/web Patch
workflow Patch
@workflow/world-testing Patch
@workflow/astro Patch
@workflow/nest Patch
@workflow/rollup Patch
@workflow/sveltekit Patch
@workflow/vite Patch
@workflow/nuxt Patch

Not sure what this means? Click here to learn what changesets are.

Click here if you're a maintainer who wants to add another changeset to this PR

@vercel

vercel Bot commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
example-nextjs-workflow-turbopack Ready Ready Preview, v0 Aug 24, 2026 10:33pm
example-nextjs-workflow-webpack Ready Ready Preview, v0 Aug 24, 2026 10:33pm
example-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workbench-astro-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workbench-express-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workbench-fastify-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workbench-hono-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workbench-nestjs-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workbench-nitro-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workbench-nuxt-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workbench-python-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workbench-sveltekit-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workbench-tanstack-start-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workbench-vite-workflow Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workflow-docs Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workflow-swc-playground Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workflow-tarballs Ready Ready Preview, v0 Aug 24, 2026 10:33pm
workflow-web Ready Ready Preview, v0 Aug 24, 2026 10:33pm

@github-actions

github-actions Bot commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

Some tests failed

❌ Failed E2E Tests

▲ Vercel Production (8 failed)

python-node (8 failed):

  • promiseAllWorkflow | wrun_41M0TYCQP90GH8HSRD6CPGTZZ9 | 🔍 observability
  • sleepingWorkflow | wrun_41M0TYDCYJ0GJ31NQ8TEC2RM0K | 🔍 observability
  • parallelSleepWorkflow | wrun_41M0TYDG8X0GYYV0DPF5TTWPQ1 | 🔍 observability
  • nullByteWorkflow | wrun_41M0TYDPF10GNAH5FBTG9D432C | 🔍 observability
  • cancelRun - cancelling a running workflow | wrun_41M0TYJFZ00GQDE241HGNYZ43C | 🔍 observability
  • cancelRun via CLI - cancelling a running workflow | wrun_41M0TYJKG60GR86KZH96M8F0ZV | 🔍 observability
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M0TYJV6W0GPABTHJT042EFNG | 🔍 observability
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M0TYKCQT0GK35M5THBBEEMBP | 🔍 observability

🌐 Cross-language Conformance (9 failed)

python (9 failed):

  • deploymentId: 'latest' is a no-op in non-Vercel worlds | wrun_01M0TYP4EWXS4MHWKXFZN1JR0F
  • promiseAllWorkflow | wrun_41M0TYCQP90GH8HSRD6CPGTZZ9
  • sleepingWorkflow | wrun_41M0TYDCYJ0GJ31NQ8TEC2RM0K
  • parallelSleepWorkflow | wrun_41M0TYDG8X0GYYV0DPF5TTWPQ1
  • nullByteWorkflow | wrun_41M0TYDPF10GNAH5FBTG9D432C
  • cancelRun - cancelling a running workflow | wrun_41M0TYJFZ00GQDE241HGNYZ43C
  • cancelRun via CLI - cancelling a running workflow | wrun_41M0TYJKG60GR86KZH96M8F0ZV
  • sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration | wrun_41M0TYJV6W0GPABTHJT042EFNG
  • resilient start: addTenWorkflow completes when run_created returns 500 | wrun_41M0TYKCQT0GK35M5THBBEEMBP

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

  • addTenWorkflow (fastify)
  • addTenWorkflow (hono)
  • addTenWorkflow (nuxt)
  • cancelRun via CLI - cancelling a running workflow (nextjs-turbopack)
  • cancelRun via CLI - cancelling a running workflow (nuxt)
  • hookClaimOnlyMutexWorkflow - hook works as a pure run mutex without payload data (nextjs-webpack)
  • hookDisposeTestWorkflow - hook token reuse after explicit disposal while workflow still running (nextjs-webpack)
  • sleepWinsRaceWorkflow (tanstack-start)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

32 infra events
  • cold-start-warmup · suite warmup (python) · at 22:32:26Z · abandoned wrun_41M0TYAYJT0GNGCGFT9CW8J31C · (+7 more)
  • run-pickup-stall · parallelSleepWorkflow (python) · at 22:32:42Z · abandoned wrun_41M0TYEKEX0GZ0RYFQZ1AGJAN9
  • run-pickup-stall · nullByteWorkflow (python) · at 22:32:42Z · abandoned wrun_41M0TYEKF10GW0JY2HGZZ2K8BR
  • run-pickup-stall · sleepingWorkflow (python) · at 22:32:42Z · abandoned wrun_41M0TYEKEX0GZ0RYFQZ1AGJAN8
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 22:32:42Z · abandoned wrun_41M0TYEKTR0GSK8P4JGC50Q9YT
  • run-pickup-stall · promiseAllWorkflow (python) · at 22:32:43Z · abandoned wrun_41M0TYEKEQ0GVKJ0FPW7DV5X07
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 22:33:15Z · abandoned wrun_41M0TYFKB30GZV9RYEXPSVWC4S
  • run-pickup-stall · sleepingWorkflow (python) · at 22:33:43Z · abandoned wrun_41M0TYGEYP0GNT5G4RBTCFB13K
  • run-pickup-stall · parallelSleepWorkflow (python) · at 22:33:43Z · abandoned wrun_41M0TYGF3N0GT1E7GFW1V9TJ21
  • run-pickup-stall · promiseAllWorkflow (python) · at 22:33:43Z · abandoned wrun_41M0TYGEYS0GY7C05YYWCVAN24
  • run-pickup-stall · nullByteWorkflow (python) · at 22:33:43Z · abandoned wrun_41M0TYGF250GVD8T6M4BGNC9MZ
  • cold-start-warmup · suite warmup (tanstack-start) · at 22:34:07Z · abandoned wrun_01M0TYGV52206SC5AEK09QTHDV
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 22:34:33Z · abandoned wrun_41M0TYGK1Q0GP2X8Z95ZQRRF08
  • cold-start-warmup · suite warmup (python) · at 22:35:18Z · abandoned wrun_01M0TYG6112GWK343TTH7VNV7D · (+7 more)
  • run-pickup-stall · promiseAllWorkflow (python) · at 22:35:33Z · abandoned wrun_01M0TYKV5PD11HWJJ5DJPTY36E
  • run-pickup-stall · sleepingWorkflow (python) · at 22:35:33Z · abandoned wrun_01M0TYKV5VB4BEN37GS66N59YK
  • run-pickup-stall · parallelSleepWorkflow (python) · at 22:35:33Z · abandoned wrun_01M0TYKV5Y18NPGYQK6XKMNBJY
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 22:35:33Z · abandoned wrun_01M0TYKV5PD11HWJJ5DJPTY36D
  • run-pickup-stall · nullByteWorkflow (python) · at 22:35:33Z · abandoned wrun_01M0TYKV60YRZTH9DG0JNDANBH
  • run-pickup-stall · promiseAllWorkflow (python) · at 22:36:33Z · abandoned wrun_01M0TYNNSJYNGNK851MQP1JB09
  • run-pickup-stall · deploymentId: 'latest' is a no-op in non-Vercel worlds (python) · at 22:36:33Z · abandoned wrun_01M0TYNNSGB77J7W5E13A3R263
  • run-pickup-stall · sleepingWorkflow (python) · at 22:36:33Z · abandoned wrun_01M0TYNNSM1WW7VT6JAS9B0VEX
  • run-pickup-stall · parallelSleepWorkflow (python) · at 22:36:33Z · abandoned wrun_01M0TYNNSQ61EMZQSZ849JNHA7
  • run-pickup-stall · nullByteWorkflow (python) · at 22:36:33Z · abandoned wrun_01M0TYNNSR4ER537QZBJ79SN2G
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 22:36:55Z · abandoned wrun_41M0TYK6V80GQ8NYJDE1PPJBXH
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 22:36:55Z · abandoned wrun_41M0TYKAY70GYG61T79VHX87WA
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 22:37:33Z · abandoned wrun_01M0TYQGD6NDH7M075CNB00CER
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 22:37:33Z · abandoned wrun_01M0TYQGDAM3PHNF2JA81T8P7Z
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 22:37:33Z · abandoned wrun_01M0TYQGDH2ZVXYM64QF0HSDRM
  • run-pickup-stall · cancelRun - cancelling a running workflow (python) · at 22:38:03Z · abandoned wrun_01M0TYRDRQC396P9TA2NFEF1YW
  • run-pickup-stall · cancelRun via CLI - cancelling a running workflow (python) · at 22:38:03Z · abandoned wrun_01M0TYRDRY07WYKZTT81X3CSEA
  • run-pickup-stall · sleepInLoopWorkflow - sleep inside loop with steps actually delays each iteration (python) · at 22:38:33Z · abandoned wrun_01M0TYSB0M52SKJN6P09C9FJ4R

E2E Test Summary

Summary
Passed Failed Skipped Total
❌ ▲ Vercel Production 3570 8 742 4320
✅ 💻 Local Development 3922 0 558 4480
✅ 📦 Local Production 3922 0 558 4480
✅ 🐘 Local Postgres 3922 0 558 4480
✅ 🪟 Windows 320 0 0 320
❌ 🌐 Cross-language Conformance 0 9 132 141
✅ vercel-http-transport 817 0 143 960
✅ vercel-multi-region 27 0 0 27
✅ vercel-ws-transport 553 0 87 640
Total 17053 17 2778 19848
Details by Category

❌ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 132 0 28
✅ astro-quickjs 132 0 28
✅ example-node 132 0 28
✅ example-quickjs 132 0 28
✅ express-node 132 0 28
✅ express-quickjs 132 0 28
✅ fastify-node 132 0 28
✅ fastify-quickjs 132 0 28
✅ hono-node 132 0 28
✅ hono-quickjs 132 0 28
✅ nest-node 132 0 28
✅ nest-quickjs 132 0 28
✅ nextjs-turbopack-node 157 0 3
✅ nextjs-turbopack-quickjs 157 0 3
✅ nextjs-webpack-node 157 0 3
✅ nextjs-webpack-quickjs 157 0 3
✅ nitro-node 132 0 28
✅ nitro-quickjs 132 0 28
✅ nuxt-node 132 0 28
✅ nuxt-quickjs 132 0 28
❌ python-node 0 8 152
✅ sveltekit-node 151 0 9
✅ sveltekit-quickjs 151 0 9
✅ tanstack-start-node 132 0 28
✅ tanstack-start-quickjs 132 0 28
✅ vite-node 132 0 28
✅ vite-quickjs 132 0 28

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 134 0 26
✅ astro-stable-quickjs 134 0 26
✅ express-stable-node 134 0 26
✅ express-stable-quickjs 134 0 26
✅ fastify-stable-node 134 0 26
✅ fastify-stable-quickjs 134 0 26
✅ hono-stable-node 134 0 26
✅ hono-stable-quickjs 134 0 26
✅ nest-stable-node 134 0 26
✅ nest-stable-quickjs 134 0 26
✅ nextjs-turbopack-canary-node 141 0 19
✅ nextjs-turbopack-canary-quickjs 141 0 19
✅ nextjs-turbopack-stable-node 160 0 0
✅ nextjs-turbopack-stable-quickjs 160 0 0
✅ nextjs-webpack-canary-node 141 0 19
✅ nextjs-webpack-canary-quickjs 141 0 19
✅ nextjs-webpack-stable-node 160 0 0
✅ nextjs-webpack-stable-quickjs 160 0 0
✅ nitro-stable-node 134 0 26
✅ nitro-stable-quickjs 134 0 26
✅ nuxt-stable-node 134 0 26
✅ nuxt-stable-quickjs 134 0 26
✅ sveltekit-stable-node 153 0 7
✅ sveltekit-stable-quickjs 153 0 7
✅ tanstack-start-node 134 0 26
✅ tanstack-start-quickjs 134 0 26
✅ vite-stable-node 134 0 26
✅ vite-stable-quickjs 134 0 26

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 134 0 26
✅ astro-stable-quickjs 134 0 26
✅ express-stable-node 134 0 26
✅ express-stable-quickjs 134 0 26
✅ fastify-stable-node 134 0 26
✅ fastify-stable-quickjs 134 0 26
✅ hono-stable-node 134 0 26
✅ hono-stable-quickjs 134 0 26
✅ nest-stable-node 134 0 26
✅ nest-stable-quickjs 134 0 26
✅ nextjs-turbopack-canary-node 141 0 19
✅ nextjs-turbopack-canary-quickjs 141 0 19
✅ nextjs-turbopack-stable-node 160 0 0
✅ nextjs-turbopack-stable-quickjs 160 0 0
✅ nextjs-webpack-canary-node 141 0 19
✅ nextjs-webpack-canary-quickjs 141 0 19
✅ nextjs-webpack-stable-node 160 0 0
✅ nextjs-webpack-stable-quickjs 160 0 0
✅ nitro-stable-node 134 0 26
✅ nitro-stable-quickjs 134 0 26
✅ nuxt-stable-node 134 0 26
✅ nuxt-stable-quickjs 134 0 26
✅ sveltekit-stable-node 153 0 7
✅ sveltekit-stable-quickjs 153 0 7
✅ tanstack-start-node 134 0 26
✅ tanstack-start-quickjs 134 0 26
✅ vite-stable-node 134 0 26
✅ vite-stable-quickjs 134 0 26

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 134 0 26
✅ astro-stable-quickjs 134 0 26
✅ express-stable-node 134 0 26
✅ express-stable-quickjs 134 0 26
✅ fastify-stable-node 134 0 26
✅ fastify-stable-quickjs 134 0 26
✅ hono-stable-node 134 0 26
✅ hono-stable-quickjs 134 0 26
✅ nest-stable-node 134 0 26
✅ nest-stable-quickjs 134 0 26
✅ nextjs-turbopack-canary-node 141 0 19
✅ nextjs-turbopack-canary-quickjs 141 0 19
✅ nextjs-turbopack-stable-node 160 0 0
✅ nextjs-turbopack-stable-quickjs 160 0 0
✅ nextjs-webpack-canary-node 141 0 19
✅ nextjs-webpack-canary-quickjs 141 0 19
✅ nextjs-webpack-stable-node 160 0 0
✅ nextjs-webpack-stable-quickjs 160 0 0
✅ nitro-stable-node 134 0 26
✅ nitro-stable-quickjs 134 0 26
✅ nuxt-stable-node 134 0 26
✅ nuxt-stable-quickjs 134 0 26
✅ sveltekit-stable-node 153 0 7
✅ sveltekit-stable-quickjs 153 0 7
✅ tanstack-start-node 134 0 26
✅ tanstack-start-quickjs 134 0 26
✅ vite-stable-node 134 0 26
✅ vite-stable-quickjs 134 0 26

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 160 0 0
✅ nextjs-turbopack-quickjs 160 0 0

❌ 🌐 Cross-language Conformance

App Passed Failed Skipped
❌ python 0 9 132

✅ vercel-http-transport

App Passed Failed Skipped
✅ example 132 0 28
✅ express 132 0 28
✅ hono 132 0 28
✅ nextjs-turbopack 157 0 3
✅ nitro 132 0 28
✅ vite 132 0 28

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

✅ vercel-ws-transport

App Passed Failed Skipped
✅ example 132 0 28
✅ express 132 0 28
✅ nextjs-turbopack 157 0 3
✅ vite 132 0 28

📋 View full workflow run

@github-actions

github-actions Bot commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit d11c0dd · Mon, 24 Aug 2026 22:55:03 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 246 (-75%) 💚 1346 🔴 (+22%) 🔻 1384 🔴 (+20%) 🔻 1484 🔴 (+18%) 🔻 30
TTFS stream 222 (-69%) 💚 1328 🔴 (+21%) 🔻 1400 🔴 (+26%) 🔻 1432 🔴 (+11%) 30
TTFS hook + stream 1264 (-8.0%) 1646 🔴 (+12%) 1722 🔴 (+14%) 2321 🔴 (+45%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 645 (-0.8%) 1882 (+10%) 2135 (+19%) 🔻 2343 (+20%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 1841 (-4.7%) 3816 (+20%) 🔻 5084 (+55%) 🔻 5570 (-34%) 💚 10
STSO 1020 steps (inline) 69 (-34%) 💚 140 (-5.4%) 164 (-2.4%) 275 (+26%) 🔻 1019
WO 1020 steps 140686 (-4.7%) 140686 (-4.7%) 140686 (-4.7%) 140686 (-4.7%) 1
CRTT first chunk (pooled) 107 (+22%) 🔻 171 (+38%) 🔻 260 (+35%) 🔻 274 (-56%) 💚 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 141 (+25%) 256 (+97%) 513 (+142%) 967 (+27%) 279 (+95%) 10
size sweep (100/s, 160B-12KB) 150 (+33%) 164 (+15%) 275 (+47%) 495 (-56%) 195 (+42%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 131 (+7%) 532 (-18%) 706 (-79%) 911 (-81%) 203 (-6%) 3
replay eve-gpt-5.6-sol-2000t (1x) 216 (+83%) 168 (+25%) 261 (+54%) 1122 (+225%) 761 (+191%) 2
replay eve-gpt-5.6-sol-2000t (2x) 150 (+22%) 207 (+13%) 274 (+4%) 691 (+25%) 268 (+18%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 147417ms → this run 140462ms (Δ -6955ms, -5%)

 50-100 ms  ┃                         main   0  this   1    +1
100-150 ms  ██████████████████████░┃  main 785  this 848   +63
150-200 ms  ███┃██                    main 216  this 143   -73
200-250 ms  ┃                         main  14  this  11    -3
250-300 ms  ┃                         main   4  this   9    +5
300-350 ms  ┃                         main   0  this   5    +5
350-400 ms  ┃                         main   0  this   1    +1
400-450 ms  ┃                         main   0  this   1    +1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+             avg         p50          p90           p99     n
control  ······▃█▃▁···  175.1 (+46%)  135 (+27%)  513 (+142%)    967 (+27%)  3000
sweep    ······▃█▂▁···  136.8 (+11%)  122 (+11%)   275 (+47%)    495 (-56%)  3000
gw 1x    ·····▁▄█▃▂···  203.3 (-39%)  119 (+18%)   706 (-79%)    911 (-81%)  5295
eve 1x   ·····▁▃█▂▁▁··  171.8 (+55%)  129 (+42%)   261 (+54%)  1122 (+225%)  5186
eve 2x   ·····▁▂█▃▁···  167.7 (+20%)  149 (+16%)    274 (+4%)    691 (+25%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  ▅█▄▆▃▄▁▆▅█  145–199ms
sweep    █▆▂▁▂▅▃▁▃▂  116–178ms
gw 1x    ██▄▄▅▁▂▄▇▅  122–275ms
eve 1x   ▂▂▃▂▁▁▄█▁▁  121–326ms
eve 2x   ▃▁▁▁▁▂▅██▂  142–224ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  ▁▅█▅▅▂▃  136–138ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  ▄▁▁▆▆▃▆▄█▅  34–56ms
sweep    ▁▃▃▃▆▄▆▄█▂  44–59ms
gw 1x    █▄▂▂▃▁▄▃▆▃  38–71ms
eve 1x   ▅▆▆▅▃▁█▅▃▃  24–39ms
eve 2x   █▇▁▃▆▅█▃▄▆  22–36ms
📜 Previous results (2)

47501c6

Fri, 21 Aug 2026 23:12:41 GMT · run logs

vercel / nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 1314 (+22%) 🔻 1396 🔴 (+21%) 🔻 1428 🔴 (+14%) 1481 🔴 (+3.6%) 30
TTFS stream 270 (-75%) 💚 1352 🔴 (+12%) 1378 🔴 (+7.1%) 1389 🔴 (-3.9%) 30
TTFS hook + stream 1534 (+45%) 🔻 1659 🔴 (+14%) 1766 🔴 (+17%) 🔻 2433 🔴 (+57%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 568 (-4.1%) 2033 (+28%) 🔻 2052 (+24%) 🔻 2111 (+27%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 1699 (+5.1%) 3860 (+28%) 🔻 4036 (+12%) 4573 (+16%) 🔻 10
STSO 1020 steps (inline) 114 (-24%) 💚 168 (-24%) 💚 189 (-27%) 💚 344 (-36%) 💚 1019
WO 1020 steps 169326 (-22%) 💚 169326 (-22%) 💚 169326 (-22%) 💚 169326 (-22%) 💚 1
CRTT first chunk (pooled) 99 (+4.2%) 152 (+2.7%) 189 (+8.6%) 376 (+62%) 🔻 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 141 (+9%) 143 (-23%) 186 (-27%) 402 (+6%) 110 (-35%) 10
size sweep (100/s, 160B-12KB) 125 (+1%) 151 (-18%) 195 (-42%) 279 (-46%) 99.5 (-25%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 152 (+26%) 189 (-74%) 414 (-87%) 757 (-84%) 466 (-86%) 3
replay eve-gpt-5.6-sol-2000t (1x) 159 (+31%) 148 (-8%) 186 (-10%) 287 (-20%) 296 (-2%) 2
replay eve-gpt-5.6-sol-2000t (2x) 135 (-13%) 200 (-38%) 268 (-73%) 380 (-83%) 246 (-38%) 3

62fcd0f

Wed, 19 Aug 2026 23:10:40 GMT · run logs

vercel / nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 1287 (+24%) 🔻 1401 🔴 (+19%) 🔻 1442 🔴 (+17%) 🔻 1732 🔴 (+4.1%) 30
TTFS stream 1260 (+25%) 🔻 1388 🔴 (+25%) 🔻 1428 🔴 (+25%) 🔻 1742 🔴 (+49%) 🔻 30
TTFS hook + stream 538 (-58%) 💚 1644 🔴 (+15%) 1690 🔴 (+15%) 1986 🔴 (+26%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 527 (-7.9%) 829 (-16%) 💚 986 (-45%) 💚 2498 (+38%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 4804 (+1.2%) 8742 (+6.0%) 9471 (-13%) 10733 (-19%) 💚 10
STSO 1020 steps (inline) 117 (-5.6%) 159 (-35%) 💚 185 (-37%) 💚 303 (-40%) 💚 1019
WO 1020 steps 158021 (-33%) 💚 158021 (-33%) 💚 158021 (-33%) 💚 158021 (-33%) 💚 1
CRTT first chunk (pooled) 101 (-4.7%) 175 (+3.6%) 372 (+84%) 🔻 3550 (+770%) 🔻 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 115 (-20%) 131 (-27%) 168 (-42%) 335 (-36%) 96.5 (-32%) 10
size sweep (100/s, 160B-12KB) 127 (-9%) 135 (-26%) 172 (-32%) 310 (-31%) 134 (-14%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 204 (+37%) 139 (-4%) 230 (+8%) 675 (-3%) 293 (+18%) 3
replay eve-gpt-5.6-sol-2000t (1x) 1837 (+1128%) 140 (-5%) 193 (±0%) 3960 (+1082%) 2212 (+683%) 2
replay eve-gpt-5.6-sol-2000t (2x) 372 (+120%) 242 (+12%) 468 (+23%) 1195 (±0%) 226 (-43%) 3
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@github-actions

github-actions Bot commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 20 1.0m ok 0
stale-read-equal-step-counts completed 14 1.0m ok 0
step-vs-step-fork completed 12 0ms ok 0
step-vs-step-fork-fenced completed 12 0ms ok 0
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision completed 17 1.0m ok 0
in-flight-before-decision-counted completed 17 1.0m ok 0
in-flight-after-decision completed 19 2.0m ok 0
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

@VaguelySerious
VaguelySerious marked this pull request as ready for review August 21, 2026 22:48
@VaguelySerious
VaguelySerious requested a review from a team as a code owner August 21, 2026 22:48
@VaguelySerious
VaguelySerious merged commit 5841558 into main Aug 25, 2026
313 of 320 checks passed
@VaguelySerious
VaguelySerious deleted the peter/windows-step-registration-cleanup branch August 25, 2026 19:07
@github-actions

Copy link
Copy Markdown
Contributor

No backport to stable for 5841558 (AI decision).

This is CI/test hardening for a failure mode that does not exist on stable: the generated __step_registrations.js that strands a deleted step fixture is main-only (on stable, step_registration appears only in the SWC plugin's Rust, and packages/next generates no such file), and the dev.test.ts changes build on main-only harness pieces (generatedStepRegistrationPath config, restoreDirectories). The tests.yml hunk likewise rewrites a Get-DevServerStatus pre-e2e probe block that stable does not have at all. Nothing here keeps stable buildable, testable, or releasable.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

584155897f75e712a1c2bc199d6d12027cd18dab

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants