Skip to content

[e2e] Expand Python conformance coverage - #3495

Merged
fantix merged 1 commit into
mainfrom
fantix/python-conformance-more
Sep 2, 2026
Merged

[e2e] Expand Python conformance coverage#3495
fantix merged 1 commit into
mainfrom
fantix/python-conformance-more

Conversation

@fantix

@fantix fantix commented Aug 12, 2026

Copy link
Copy Markdown
Member

Summary

Expands the Python workbench from a protocol smoke test into a cross-language conformance target driven by the existing TypeScript e2e suite.

  • Ports 48 Python fixtures covering concurrency, sleeps, retries, errors, streams, hooks, child runs, metadata, and run attributes.
  • Adds language-aware fixture and capability gating while keeping the shared test driver and assertions as the source of truth.
  • Uses vercel-workflow directly, with the resolved vercel-py revision recorded in uv.lock.
  • Simplifies the ASGI entrypoint and relies on the Vercel Python framework integration without a legacy builds override or the umbrella vercel package.
  • Makes a small set of otherwise portable assertions language-neutral.

Why

Running the same driver against TypeScript and Python tests the actual cross-language contract: serialized inputs and outputs, durable event ordering, queue delivery, retries, hooks, streams, and CLI hydration. The conformance manifest also acts as a ratchet: once a Python fixture is enabled, removing it or regressing it fails the suite instead of silently skipping coverage.

Coverage

The enabled fixtures exercise:

  • positional arguments, parallel steps, races, durable sleeps, cancellation, and resilient start;
  • binary and structured streams, including writable handles forwarded to child runs;
  • retry policy and serialized error identity and cause chains across step and workflow boundaries;
  • hook metadata, repeated payloads, disposal, token reuse, conflicts, owner discovery, mutex, adoption, and forwarding patterns;
  • child workflow creation and result retrieval;
  • caught and uncaught unregistered-step failures;
  • workflow and step attributes, including persistence when a run fails.

Known limitations

Two enabled tests have explicit capability exemptions in e2e-conformance.json:

  • the final fire-and-forget attribute write does not begin before the Python workflow coroutine returns;
  • reserved-attribute validation rejects the input correctly, but its message does not name the opt-in escape hatch expected by the shared assertion.

Fixtures that require Python APIs or failure semantics not implemented yet remain outside the enabled fixture list.

Validation

  • Local Python conformance: 64 passed | 77 skipped (141)
  • Vercel Python conformance: 62 passed | 98 skipped (160)
  • pnpm build
  • uv lock --check
  • Preview deployment using only vercel-workflow

@fantix
fantix requested a review from a team as a code owner August 12, 2026 18:09
@changeset-bot

changeset-bot Bot commented Aug 12, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 7881413

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

This PR includes no changesets

When changesets are added to this PR, you'll see the packages that this PR includes changesets for and the associated semver types

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@vercel

vercel Bot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
example-nextjs-workflow-turbopack Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
example-nextjs-workflow-webpack Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
example-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workbench-astro-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workbench-express-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workbench-fastify-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workbench-hono-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workbench-nestjs-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workbench-nitro-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workbench-nuxt-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workbench-python-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workbench-sveltekit-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workbench-tanstack-start-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workbench-vite-workflow Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workflow-swc-playground Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workflow-tarballs Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
workflow-web Ready Ready Preview, v0 Sep 1, 2026 3:12pm UTC
1 Skipped Deployment
Project Deployment Actions Updated
workflow-docs Skipped Skipped v0 Sep 1, 2026 3:12pm UTC

@github-actions

github-actions Bot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

🧪 E2E Test Results

All tests passed

⚠️ Flaky E2E Tests (passed on retry)

These tests failed at least once and passed on a retry. A recurring entry here is a real race worth investigating.

10 flaky tests
  • an sfo1 reader sees chunks of an IN-PROGRESS iad1 stream (nextjs-turbopack)
  • cancelRun via CLI - cancelling a running workflow (express)
  • cancelRun via CLI - cancelling a running workflow (hono)
  • FatalError fails immediately without retries (astro)
  • outputStreamInsideStepWorkflow - getWritable() called inside step functions (vite)
  • positive startIndex (skips first chunk) (vite)
  • promiseAllWorkflow (nitro)
  • sleepWinsRaceWorkflow (vite)
  • workflow throw of a non-Error value round-trips verbatim as cause (nitro)
  • workflow throw of a non-Error value round-trips verbatim as cause (tanstack-start)

🛠 Infra Events (absorbed by the harness)

Platform anomalies the e2e harness detected and worked around (e.g. a run the queue never picked up, replaced by a fresh run). Clustered timestamps indicate a backend blip; a steady drip indicates a platform issue worth escalating.

  • cold-start-warmup · suite warmup (tanstack-start) · at 15:12:49Z · abandoned wrun_01M1EREPCY6PXC0YWAYEJHYAEK
  • run-pickup-stall · hookCleanupTestWorkflow - hook token reuse after workflow completion (nextjs-webpack) · at 15:17:31Z · abandoned wrun_01M1ERQH9T1B99QVKH246FRZ61
  • run-pickup-stall · concurrent hook token conflict - two workflows cannot use the same hook token simultaneously (nextjs-webpack) · at 15:17:32Z · abandoned wrun_01M1ERQJGYJ9QETDS5BMXB39DE

E2E Test Summary

Summary
Passed Failed Skipped Total
✅ ▲ Vercel Production 3632 0 688 4320
✅ 💻 Local Development 3922 0 558 4480
✅ 📦 Local Production 3922 0 558 4480
✅ 🐘 Local Postgres 3922 0 558 4480
✅ 🪟 Windows 320 0 0 320
✅ 🌐 Cross-language Conformance 64 0 77 141
✅ vercel-http-transport 817 0 143 960
✅ vercel-multi-region 27 0 0 27
✅ vercel-ws-transport 553 0 87 640
Total 17179 0 2669 19848
Details by Category

✅ ▲ Vercel Production

App Passed Failed Skipped
✅ astro-node 132 0 28
✅ astro-quickjs 132 0 28
✅ example-node 132 0 28
✅ example-quickjs 132 0 28
✅ express-node 132 0 28
✅ express-quickjs 132 0 28
✅ fastify-node 132 0 28
✅ fastify-quickjs 132 0 28
✅ hono-node 132 0 28
✅ hono-quickjs 132 0 28
✅ nest-node 132 0 28
✅ nest-quickjs 132 0 28
✅ nextjs-turbopack-node 157 0 3
✅ nextjs-turbopack-quickjs 157 0 3
✅ nextjs-webpack-node 157 0 3
✅ nextjs-webpack-quickjs 157 0 3
✅ nitro-node 132 0 28
✅ nitro-quickjs 132 0 28
✅ nuxt-node 132 0 28
✅ nuxt-quickjs 132 0 28
✅ python-node 62 0 98
✅ sveltekit-node 151 0 9
✅ sveltekit-quickjs 151 0 9
✅ tanstack-start-node 132 0 28
✅ tanstack-start-quickjs 132 0 28
✅ vite-node 132 0 28
✅ vite-quickjs 132 0 28

✅ 💻 Local Development

App Passed Failed Skipped
✅ astro-stable-node 134 0 26
✅ astro-stable-quickjs 134 0 26
✅ express-stable-node 134 0 26
✅ express-stable-quickjs 134 0 26
✅ fastify-stable-node 134 0 26
✅ fastify-stable-quickjs 134 0 26
✅ hono-stable-node 134 0 26
✅ hono-stable-quickjs 134 0 26
✅ nest-stable-node 134 0 26
✅ nest-stable-quickjs 134 0 26
✅ nextjs-turbopack-canary-node 141 0 19
✅ nextjs-turbopack-canary-quickjs 141 0 19
✅ nextjs-turbopack-stable-node 160 0 0
✅ nextjs-turbopack-stable-quickjs 160 0 0
✅ nextjs-webpack-canary-node 141 0 19
✅ nextjs-webpack-canary-quickjs 141 0 19
✅ nextjs-webpack-stable-node 160 0 0
✅ nextjs-webpack-stable-quickjs 160 0 0
✅ nitro-stable-node 134 0 26
✅ nitro-stable-quickjs 134 0 26
✅ nuxt-stable-node 134 0 26
✅ nuxt-stable-quickjs 134 0 26
✅ sveltekit-stable-node 153 0 7
✅ sveltekit-stable-quickjs 153 0 7
✅ tanstack-start-node 134 0 26
✅ tanstack-start-quickjs 134 0 26
✅ vite-stable-node 134 0 26
✅ vite-stable-quickjs 134 0 26

✅ 📦 Local Production

App Passed Failed Skipped
✅ astro-stable-node 134 0 26
✅ astro-stable-quickjs 134 0 26
✅ express-stable-node 134 0 26
✅ express-stable-quickjs 134 0 26
✅ fastify-stable-node 134 0 26
✅ fastify-stable-quickjs 134 0 26
✅ hono-stable-node 134 0 26
✅ hono-stable-quickjs 134 0 26
✅ nest-stable-node 134 0 26
✅ nest-stable-quickjs 134 0 26
✅ nextjs-turbopack-canary-node 141 0 19
✅ nextjs-turbopack-canary-quickjs 141 0 19
✅ nextjs-turbopack-stable-node 160 0 0
✅ nextjs-turbopack-stable-quickjs 160 0 0
✅ nextjs-webpack-canary-node 141 0 19
✅ nextjs-webpack-canary-quickjs 141 0 19
✅ nextjs-webpack-stable-node 160 0 0
✅ nextjs-webpack-stable-quickjs 160 0 0
✅ nitro-stable-node 134 0 26
✅ nitro-stable-quickjs 134 0 26
✅ nuxt-stable-node 134 0 26
✅ nuxt-stable-quickjs 134 0 26
✅ sveltekit-stable-node 153 0 7
✅ sveltekit-stable-quickjs 153 0 7
✅ tanstack-start-node 134 0 26
✅ tanstack-start-quickjs 134 0 26
✅ vite-stable-node 134 0 26
✅ vite-stable-quickjs 134 0 26

✅ 🐘 Local Postgres

App Passed Failed Skipped
✅ astro-stable-node 134 0 26
✅ astro-stable-quickjs 134 0 26
✅ express-stable-node 134 0 26
✅ express-stable-quickjs 134 0 26
✅ fastify-stable-node 134 0 26
✅ fastify-stable-quickjs 134 0 26
✅ hono-stable-node 134 0 26
✅ hono-stable-quickjs 134 0 26
✅ nest-stable-node 134 0 26
✅ nest-stable-quickjs 134 0 26
✅ nextjs-turbopack-canary-node 141 0 19
✅ nextjs-turbopack-canary-quickjs 141 0 19
✅ nextjs-turbopack-stable-node 160 0 0
✅ nextjs-turbopack-stable-quickjs 160 0 0
✅ nextjs-webpack-canary-node 141 0 19
✅ nextjs-webpack-canary-quickjs 141 0 19
✅ nextjs-webpack-stable-node 160 0 0
✅ nextjs-webpack-stable-quickjs 160 0 0
✅ nitro-stable-node 134 0 26
✅ nitro-stable-quickjs 134 0 26
✅ nuxt-stable-node 134 0 26
✅ nuxt-stable-quickjs 134 0 26
✅ sveltekit-stable-node 153 0 7
✅ sveltekit-stable-quickjs 153 0 7
✅ tanstack-start-node 134 0 26
✅ tanstack-start-quickjs 134 0 26
✅ vite-stable-node 134 0 26
✅ vite-stable-quickjs 134 0 26

✅ 🪟 Windows

App Passed Failed Skipped
✅ nextjs-turbopack-node 160 0 0
✅ nextjs-turbopack-quickjs 160 0 0

✅ 🌐 Cross-language Conformance

App Passed Failed Skipped
✅ python 64 0 77

✅ vercel-http-transport

App Passed Failed Skipped
✅ example 132 0 28
✅ express 132 0 28
✅ hono 132 0 28
✅ nextjs-turbopack 157 0 3
✅ nitro 132 0 28
✅ vite 132 0 28

✅ vercel-multi-region

App Passed Failed Skipped
✅ nextjs-turbopack 27 0 0

✅ vercel-ws-transport

App Passed Failed Skipped
✅ example 132 0 28
✅ express 132 0 28
✅ nextjs-turbopack 157 0 3
✅ vite 132 0 28

📋 View full workflow run

@github-actions

github-actions Bot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

Sim World

Simulated world deterministic testing for races. Traces

🟠 world-sim scenario book — 1 fail of 41 total

fence=per-spec

scenario outcome events virt replay violations
smoke-no-steps completed 3 0ms ok 0
smoke-one-step completed 6 0ms ok 0
hook-at-step-started completed 12 0ms ok 0
hook-at-step-completed completed 12 0ms ok 0
hook-at-hook-created completed 12 0ms ok 0
deadline-hook-wins completed 7 1.0h ok 0
deadline-expires completed 7 1.0h ok 0
long-sleep completed 11 30.0d ok 0
hook-never-arrives stalled 3 0ms skipped 0
step-retries-twice completed 10 2.0s ok 0
parallel-steps completed 9 0ms ok 0
hook-on-execution-state completed 12 0ms ok 0
peek-hook-before-branch completed 12 0ms ok 0
peek-hook-after-branch completed 12 0ms ok 0
peek-hook-at-registration completed 12 0ms ok 0
race-hook-before-probe completed 12 0ms ok 0
race-hook-after-probe completed 12 0ms ok 0
race-duplicate-delivery completed 13 0ms ok 0
attr-hook-before-step completed 11 0ms ok 0
attr-hook-after-step completed 11 0ms ok 0
attr-from-step-body completed 13 0ms ok 0
fork-hook-after-timeout completed 14 1.0m ok 0
fork-hook-before-timeout completed 14 1.0m ok 0
count-hook-after-timeout completed 17 1.0m ok 0
count-hook-before-timeout completed 20 1.0m ok 0
stale-read-step-count-fork completed 20 1.0m ok 0
stale-read-equal-step-counts completed 14 1.0m ok 0
step-vs-step-fork completed 12 0ms ok 0
step-vs-step-fork-fenced completed 12 0ms ok 0
fence-catches-benign-direction completed 12 5ms ok 0
in-flight-before-decision completed 17 1.0m ok 0
in-flight-before-decision-counted completed 17 1.0m ok 0
in-flight-after-decision completed 19 2.0m ok 0
stale-read-step-count-fork-fenced completed 20 1.0m ok 0
fork-hook-wins completed 13 1.0m ok 0
fork-timeout-wins completed 13 1.0m ok 0
unclaimed-payload-under-fork completed 17 1.0m ok 0
claimed-payload-under-fork completed 17 1.0m ok 0
writers-independent-step-bodies completed 12 0ms ok 0
writers-scripted-tempo completed 12 0ms ok 0
cancel-mid-step cancelled 7 0ms skipped 0

Full trace: world-sim.txt

@github-actions

github-actions Bot commented Aug 12, 2026

Copy link
Copy Markdown
Contributor

📊 Workflow Benchmarks

commit 7881413 · Tue, 01 Sep 2026 15:32:38 GMT · run logs

Backend: vercel · app: nextjs-turbopack

Metric Scenario Best (ms) P75 (ms) P90 (ms) P99 (ms) Samples
TTFS step 220 (+20%) 🔻 1370 🔴 (+29%) 🔻 1428 🔴 (+28%) 🔻 1768 🔴 (+24%) 🔻 30
TTFS stream 228 (+6.0%) 1296 🔴 (+19%) 🔻 1406 🔴 (+28%) 🔻 1471 🔴 (+26%) 🔻 30
TTFS hook + stream 1469 (+20%) 🔻 1565 🔴 (+17%) 🔻 1621 🔴 (+17%) 🔻 1798 🔴 (+15%) 🔻 30
Fan-out TTFS Promise.all(100 steps) 599 (±0%) 2023 (+27%) 🔻 2105 (+20%) 🔻 2486 (+32%) 🔻 10
Fan-out TTLS Promise.all(100 steps) 1862 (+12%) 3730 (+12%) 3805 (+7.2%) 8061 (+21%) 🔻 10
STSO 1020 steps (inline) 125 (+13%) 160 (+8.1%) 185 (+6.3%) 257 (-27%) 💚 1019
WO 1020 steps 160029 (+7.4%) 160029 (+7.4%) 160029 (+7.4%) 160029 (+7.4%) 1
CRTT first chunk (pooled) 98 (+1.0%) 133 (-11%) 172 (±0%) 285 (+11%) 28

Streams

Scenario CRTT 1st p75 p90 p99 CDV max iters
paced control (100/s, 60B) 119 (-18%) 154 (-14%) 344 (-12%) 566 (-24%) 137 (-29%) 10
size sweep (100/s, 160B-12KB) 108 (-16%) 154 (-17%) 219 (-35%) 382 (-19%) 158 (+3%) 10
replay gateway-gpt-5.4-nano-2000t (1x) 111 (-16%) 129 (-19%) 181 (-28%) 358 (-40%) 271 (-18%) 3
replay eve-gpt-5.6-sol-2000t (1x) 138 (-18%) 134 (-21%) 180 (-24%) 318 (-63%) 244 (-60%) 2
replay eve-gpt-5.6-sol-2000t (2x) 146 (+15%) 184 (-23%) 255 (-28%) 416 (-35%) 275 (-3%) 3
📈 STSO distribution vs main (inline / queue-hop histograms)

1020 steps (inline)

Cumulative STSO time: main 148790ms → this run 158789ms (Δ +9999ms, +7%)

100-150 ms  ████████████████┃███████  main 779  this 556  -223
150-200 ms  ██████░░░░░░┃             main 181  this 416  +235
200-250 ms  ┃                         main  31  this  35    +4
250-300 ms  ┃                         main  11  this   8    -3
300-350 ms  ┃                         main   6  this   2    -4
350-400 ms  ┃                         main   6  this   1    -5
400-450 ms  ┃                         main   1  this   1    +0
450-500 ms  ┃                         main   1  this   0    -1
500-550 ms  ┃                         main   1  this   0    -1
600-650 ms  ┃                         main   1  this   0    -1
650-700 ms  ┃                         main   1  this   0    -1
📈 CRTT drill-down vs main (RTT distributions & profiles)
variant  RTT 1ms→5s+             avg         p50         p90         p99     n
control  ······▄█▁▁···    129 (-15%)   115 (-9%)  344 (-12%)  566 (-24%)  3000
sweep    ······▄█▁····  126.7 (-16%)   117 (-7%)  219 (-35%)  382 (-19%)  3000
gw 1x    ·····▁▆█▁····  115.6 (-16%)  106 (-12%)  181 (-28%)  358 (-40%)  5295
eve 1x   ·····▁▆█▁····  116.5 (-27%)  106 (-17%)  180 (-24%)  318 (-63%)  5186
eve 2x   ·····▁▂█▃····  154.7 (-18%)  138 (-16%)  255 (-28%)  416 (-35%)  7779

RTT over stream progress (avg per tenth of stream, bars scaled min→max):

control  ▄▄▄▃▁▂██▄▂  111–154ms
sweep    ██▄▂▂▃▂▁▄▄  109–157ms
gw 1x    █▃▅▅▆▃▁▃▂█  98–133ms
eve 1x   ▅▁▂▂▃▄▅█▄▁  101–141ms
eve 2x   ▁▁▂▂▁▂▄█▃▃  134–220ms

RTT by chunk size (avg per log size bin, ~160B → ~12KB serialized, bars scaled min→max):

sweep  ▆██▅▃▂▁  124–129ms

Delivery jitter over stream progress (avg positive CDV per tenth of stream, bars scaled min→max):

control  ▁▄▅▄▂█▄▄▂▅  27–46ms
sweep    ▁▅▅▁▅█▇▂▆▃  45–54ms
gw 1x    ▅▄▆▃▆▃▂▄▁█  29–39ms
eve 1x   █▃▄▂▇▆▇█▇▁  21–26ms
eve 2x   ▄▄▁▅▃▃▃▁▁█  23–29ms
ℹ️ Metric definitions & methodology

Streams: first-chunk RTT (the stream-open path, before any buffering/backpressure), CRTT percentiles, and worst delivery stall (CDV max). Cells are medians across iterations; per-run values in the artifacts. No 🔴/🟢 marks until targets attach.

The collapsed STSO distribution section above buckets every step gap, split inline (same warm process — pure framework overhead) vs queue-hop (fresh process — dispatch, reinit, replay). = main, = this run, = fill.

The collapsed CRTT drill-down: per-variant RTT histograms (fixed log bins, · = empty) and mean RTT/positive-CDV profile lines over stream progress and chunk size. Histograms, avgs, and profiles merge exactly across runs; p50–p99 are percentile-of-percentiles. Per-index rows live in the artifacts.

Best/P75/P90/P99 deltas compare against the most recent benchmark run on main at the time of this run. 🔻 flags a delta worse than +15%, 💚 one better than −15%.

Metrics — TTFS: time to first step body (in-deployment start() → first step body) · Fan-out TTFS: fan-out time to first step (in-deployment start() → first of the parallel step bodies to complete) · Fan-out TTLS: fan-out time to last step (in-deployment start() → last of the parallel step bodies to complete, i.e. when the Promise.all resolves) · STSO: step-to-step overhead (gap between consecutive step bodies) · WO: workflow overhead (whole-run time outside step bodies, in-deployment anchored) · CRTT: chunk round-trip time (per-chunk write → read latency, one clock domain: deployment → stream backend → same deployment) · CDV: chunk delay variation / delivery jitter (inter-arrival gap minus inter-write gap per seq-adjacent pair; skew-free; the row is each run's MAX positive value, so one stall moves it)

Scenarios — step: one trivial no-op step, no stream; no hooks, so the run stays in turbo mode (in-process fast path) · stream: one streaming step; no hooks, so the run stays in turbo mode (in-process fast path) · hook + stream: registers a hook before one step, which exits turbo mode (dispatch path) · 1020 steps: 1020 trivial sequential steps; STSO is measured between consecutive steps in the given step ranges, and WO is the whole-run overhead outside step bodies · Promise.all(100 steps): 100 trivial no-op steps started together in a single Promise.all; Fan-out TTFS is the first of them to complete and Fan-out TTLS the last, both from the in-deployment clientStart, so their gap is the spread the runtime adds across the fan-out · paced control (100/s, 60B): the control: 300 tiny (~60B) deltas metronome-paced at 100/s — zero workload structure, so it reads the transport floor and flush cadence, and disambiguates transport-wide vs workload-specific when a replay row moves · size sweep (100/s, 160B-12KB): same pacing as the control with deltas padded in rotation across seven log-spaced sizes (~160B–12KB) — rotation decouples size from stream position, so it isolates whether chunk size causes latency · replay gateway-gpt-5.4-nano-2000t (1x): raw provider SSE cadence captured at the AI gateway boundary (gpt-5.4-nano, the most popular gateway model; per-token deltas p50 208B = the modal production chunk size), replayed exactly as measured — the typical customer's workload; its CDV is the typical customer's real delivery jitter · replay eve-gpt-5.6-sol-2000t (1x): a captured eve turn (gpt-5.6-sol, the most-used demanding eve model; ~2000 output tokens = production p50 turn length) replayed exactly as measured — eve's envelope protocol re-ships the cumulative message so sizes ramp 142B→13KB; the demanding outlier tenant's reality · replay eve-gpt-5.6-sol-2000t (2x): the same eve capture at 2x — the headroom/stress row; real fast-tier models emit the same chunk sizes at proportionally higher rate, so time compression is a faithful speed model · first chunk (pooled): every run's seq-0 RTT pooled across all stream scenarios — the first chunk precedes any workload differentiation, so pooling samples one shared stream-open path with exact percentiles

Replay cadences (semantic sha256) — eve-gpt-5.6-sol-2000t eaf22f5946e7c61f3c65c7006d550df180cfabd4e706254a09f22aec0cfb420d · gateway-gpt-5.4-nano-2000t 6f24ac518b6b83ff1d0e85a5fe78230db192716d66a7fc6b2fe022752001d041

🔴 marks a percentile over its target (within target is left unmarked). Targets (p75/p90/p99, ms) — TTFS 200/300/600

All timestamps are deployment-side; runs are triggered in-deployment, so the CI runner and api.vercel.com sit outside every measured window. TTFS = start() → first step body (includes dispatch + any cold start); Fan-out TTFS/TTLS = first/last step completion of one Promise.all from the same anchor (the gap is the runtime’s fan-out spread); STSO/WO between step bodies; CRTT inside the workflow (excludes the api.vercel.com read path).

Cold starts stay in the numbers (real bursty-workload latency, inflates P75+); Best is the warm floor.

@vercel
vercel Bot temporarily deployed to Preview – workflow-docs August 12, 2026 20:31 Inactive
@socket-security

socket-security Bot commented Aug 27, 2026

Copy link
Copy Markdown

Review the following changes in direct dependencies. Learn more about Socket for GitHub.

Diff Package Supply Chain
Security
Vulnerability Quality Maintenance License
Updatedpypi/​uvicorn@​0.52.3 ⏵ 0.52.498 +1100100100100

View full report

@github-actions

github-actions Bot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor
Framework Flow route Step reg. Framework output
hono 200.9 KiB (±0) 40.6 KiB (±0) 1.77 MiB (±0)
nextjs-turbopack 206.3 KiB (±0) 439 B (±0) 765.6 KiB (±0)
About these numbers

Sizes are gzip; parentheses show the change against main.
Flow route and Step reg. gate this job, on raw bytes rather than the gzip shown, at max(2%, 50.0 KiB). Framework output is informational.

7881413 · run

@fantix

fantix commented Sep 1, 2026

Copy link
Copy Markdown
Member Author

@vercel/workflow this PR is ready

@github-actions

github-actions Bot commented Sep 2, 2026

Copy link
Copy Markdown
Contributor

No backport to stable for 144b6d7 (AI decision).

This is test-coverage expansion, not a stability fix: it ports 48 new Python fixtures, adds language-aware capability gating, and switches the Python workbench's dependency to track vercel-py's main branch. Almost all of it lands in workbench/python, which I verified does not exist on origin/stable (git ls-tree origin/stable -- workbench/python returns nothing), and the packages/core/e2e/e2e.test.ts edits only exist to make shared assertions language-neutral for that absent app. The commit message notes it also fixes specVersion 7 breaking the Python e2e lane, but that fix lives in the vercel-py revision bump for an app stable does not have, so there is nothing to repair there.

To override, re-run the Backport to stable workflow manually via workflow_dispatch and paste this commit SHA into the ref input:

144b6d760188866958bbd32bc06f0b5ac5e80b85

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants