Skip to content

fix: keep retryable and 5xx status when sanitizing internal errors - #893

Draft
posthog[bot] wants to merge 1 commit into
mainfrom
posthog-self-driving/fixapi-keep-the-retryable-503-status-35d886
Draft

posthog[bot] wants to merge 1 commit into
mainfrom
posthog-self-driving/fixapi-keep-the-retryable-503-status-35d886

Conversation

@posthog

@posthog posthog Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Problem

  • Platform API callers get a non-retryable 500 when the database pooler drops a connection mid-query. The intended response is a retryable 503. So the manager runtime does not retry a short fault.
  • The API takes its status from AlienError.toExternal(). That function changes every internal error to GENERIC_ERROR, retryable: false, status 500.
  • The same thing happens to all retryable internal errors (for example the pool checkout timeout). The platform-side 503 classification has no effect until this is fixed.

Origin

  • Error tracking: issue 1, issue 2
  • First signal: 2026-09-30
  • Inbox report: open
  • Task started by: auto-start, after the report was rated P3 and ready to fix

Changes

  • An internal error keeps retryable and its 5xx status. The code, message, context, hint and source stay hidden.
  • A non-5xx status on an internal error becomes 500. An internal error is a server fault, so a 4xx (for example an upstream 401 wrapped by AlienError.from(Response)) must not blame the caller.
Internal error Before After
retryable, 503 500, not retryable 503, retryable
not retryable, 500 500, not retryable 500, not retryable (no change)
upstream 401 500, not retryable 500, not retryable (no change)
  • Tests: a connection-lost error (internal, retryable, 503) comes out as a retryable 503. An internal 401 comes out as a 500.

Note

The external response now shows whether an internal fault is retryable and which 5xx class it has. It shows no message or context.

Testing

  • pnpm test in packages/core: 101 passed.
  • pnpm test:ts and pnpm format-and-lint pass. The one Biome warning was there before this change.

Agent context

  • Rust AlienError::into_external (crates/alien-error) has the same behavior. I did not change it here: no Rust HTTP server calls into_external_response, and the CLI and worker runtime use into_external only for display. It can follow in a separate PR for parity.
  • The report also asks to keep the SQLSTATE out of the error tracking grouping text. That code is in the platform repository, not in this one.
  • The platform API gets this fix when it updates @alienplatform/core.

Created with PostHog Desktop from this inbox report.

🤖 Generated with Claude Code

toExternal() changed every internal error to a non-retryable 500. A
retryable internal 503 (for example a dropped database connection)
reached API callers as a 500 that they must not retry.

The sanitized error now keeps `retryable` and a 5xx status. It still
hides the code, message, context, hint and source. A non-5xx status
becomes 500, because a 4xx would blame the caller for a server fault.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

Generated-By: PostHog Desktop
Task-Id: a9e4e380-d99f-4011-bb75-d4e5ae320c26
@enclave-ai

enclave-ai Bot commented Oct 5, 2026

Copy link
Copy Markdown

Enclave skipped this draft pull request. It will run automatically when you mark the PR as Ready for review.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

0 participants