Skip to content

fix(cloud): a chat whose machine E2B can't start says so and keeps retrying - #181

Merged
andrewcai8 merged 3 commits into
mainfrom
fix/wake-failure-shown
Oct 6, 2026
Merged

andrewcai8 merged 3 commits into
mainfrom
fix/wake-failure-shown

Conversation

@andrewcai8

Copy link
Copy Markdown
Owner

On 2026-10-06 E2B could not resume one chat's paused sandbox (504 "Failed to place sandbox: placement timed out") for over an hour; the chat showed "updating" the whole time. The placement failure came back as a plain unknown refusal, the host then tried an upgrade to recover the box (reported as updating), and the client showed that machine state over the refusal.

  • Resume refusals caused by the provider (E2B 504 placement, 503 capacity, other 5xx) carry cause: "provider-unavailable", and the host skips the upgrade attempt for them. Presence answers carry providerUnavailableAt, cleared on the next other answer. Both fields are optional and lenient.
  • Clients show a new unavailable status: "E2B couldn't start this chat's cloud machine yet. The problem is on their side. Retrying on its own, last tried at …" on web (banner, sidebar pill) and mobile (notice, status, row), and re-poll every 15 s while it lasts.

Not classified by E2B's error_code: e2b@2.49.0 drops it from SandboxError, so status is the signal. Namespace capacity errors aren't classified yet. Tests: a real 504 through manager.resume yields the provider-unavailable refusal; 502 too; the upgrade is skipped; presence reports and clears the timestamp; client status, banner copy and polling.

🤖 Generated with Claude Code

@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Oct 6, 2026
@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire 5.0 KiB 5.0 KiB 0 B (0.0%) 6.8 KiB ✅
Codex Thread snapshot wire 3.8 KiB 3.8 KiB 0 B (0.0%) 4.9 KiB ✅
Codex Live turn WebSocket wire 1.2 KiB 1.2 KiB 0 B (0.0%) 2.0 KiB ✅
Codex Live turn WebSocket decoded 20.9 KiB 20.9 KiB 0 B (0.0%) 29.3 KiB ✅
Codex Live turn messages 2 2 0 (0.0%) 8 ✅
Claude Total thread wire 5.0 KiB 5.0 KiB 0 B (0.0%) 6.8 KiB ✅
Claude Thread snapshot wire 3.8 KiB 3.8 KiB 0 B (0.0%) 4.9 KiB ✅
Claude Live turn WebSocket wire 1.2 KiB 1.2 KiB 0 B (0.0%) 2.0 KiB ✅
Claude Live turn WebSocket decoded 21.2 KiB 21.2 KiB 0 B (0.0%) 29.3 KiB ✅
Claude Live turn messages 2 2 0 (0.0%) 8 ✅

Baseline: f829813 · PR result: 0b67461 · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 108.5 KiB
  • Claude decoded thread snapshot: 108.8 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

andrewcai8 and others added 3 commits October 6, 2026 18:36
…trying

When E2B could not place a paused sandbox (504 placement timeout, 503 no
capacity, other 5xx), the host refused the resume as `unknown`, then ran the
upgrade that recovers a stuck guest, which the host reports as "updating".
Clients showed "waking" or "updating" for as long as E2B stayed down.

The refusal now carries `cause: "provider-unavailable"`, which skips the
pointless upgrade. The host remembers when each machine's provider last failed
and its presence answer carries `providerUnavailableAt`, so every client and
the host's own wake-ahead share it. Web and mobile banners, rows and the
floating status say "E2B couldn't start this chat's cloud machine yet", that
the problem is on their side, and when it last tried. Both fields are optional
and decode leniently, so older clients and hosts are unaffected. The host's
and the client's retry backoff are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ies it

The host reported when a machine's provider last failed until another resume
answered, so a settled or archived chat the user had opened and left kept an
amber "Retrying" label and 15-second polling with nothing retrying. Presence
now reports the failure only while a wake is in flight or wake-ahead would
retry the box, and for no longer than the longest backoff.

A timeout no longer claims the problem is on E2B's side. The refusal and the
presence answer carry `provider-unavailable` (E2B answered with a 5xx) or
`provider-unreachable` (it did not answer in time), and the client says
"Couldn't reach E2B yet" for the second.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…peScript can load

The memory fixture and typecheck rejected the constructor parameter property.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@andrewcai8
andrewcai8 force-pushed the fix/wake-failure-shown branch from 735b966 to 0b67461 Compare October 6, 2026 17:37
@andrewcai8
andrewcai8 merged commit f683eb6 into main Oct 6, 2026
25 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant