Repository navigation
Colliding migration IDs silently skip migrations; one undecodable event row bricks startup #8896
Description
Activity
Possible dupe of #8789.
Independent recurrence of mechanism 2 (migrator half), this time on official nightly 0.0.39 with no fork build involved in the skipping step.
Ledger vs code (production DB, journal timestamps UTC):
ID name in DB name in 0.0.39 42 ProjectionThreadLinkedPullRequest ProjectionThreadLinkedPullRequest 43 ProjectionThreadsUnsettledAt ProjectionThreadsUnsettledAt 44 AuthSessionClientInstanceId ClearAutomaticProjectModelDefaults 45 ProjectionTasks ProjectionProjectsAutoPull 46 CleanOscPollution RepairAutomaticSettlementTimestamps IDs 44–46 in the journal came from an earlier fork line (my own dev branch, applied 2026-08-25); upstream later reused the same slots. Max recorded ID was already 46, so 044/045/046 never ran — zero migration log output — and the backend child crash-looped (exit 1, ~35s period) on:
PersistenceSqlError ... ProjectionSnapshotQuery.getCommandReadModel:listProjects ... [cause]: Error: no such column: auto_pullwhile the desktop window never became ready. I repaired by hand-applying the three skipped migrations (045 is a guarded ALTER, 044/046 are idempotent UPDATEs) after snapshotting the DB; backend has been stable since, health endpoint reporting 0.0.39.
Two notes for the fix discussion:
- The skipped migration here is from feat(projects): automatically pull clean default branches #9277 (merged Sep 2), so this failure mode is reachable by any user whose journal hit 46 through another line — not just fork users.
- A fatal preflight would have bricked this database outright (44–46 legitimately differ). Repair-by-new-ID plus a loud non-fatal warning seems the safer shape.
PR with both: #9312 (repair migration 047 + slot-collision warning in
runMigrations, with a regression test that encodes exactly this journal state).- added a commit that references this issue
on Sep 15, 2026 Hit the same defect on Windows, and I think this case closes the "but you ran a fork build" escape hatch: no fork was involved, and the trigger was a stock upgrade to current stable v0.0.42.
Reporting via
t3 triage.What happened
Desktop app offered an update, I took it, clicked the restart button, and the ADE never came back. Reinstalling and rebooting the machine changed nothing. The backend crash-looped every ~40s with the error from this issue:
ERROR (#3): PersistenceDecodeError: Decode error in OrchestrationEventStore.readFromSequence:decodeRows: Composite(Pointer(Composite(Pointer(Composite(Pointer(AnyOf(Composite(Pointer(AnyOf(AnyOf())))))))))) [cause]: SchemaError: Expected "web" | "desktop" | "mobile" | "cli" at [25]["metadata"]["origin"]["surface"]Two things this adds to the report above
1. v0.0.42 turns any pre-existing undecodable row into a boot failure, on the supported upgrade path.
v0.0.42 adds a projector that does not exist in v0.0.37:
projection.attachment-cleanup. It arrives with a fresh watermark of0, andProjectionPipeline.ts:2124then drags the read back to the very beginning regardless of how far everything else has advanced:const cleanupStart = Math.min( cleanupState?.lastAppliedSequence ?? 0, ...projectors.map((projector) => byProjector.get(projector.name)?.lastAppliedSequence ?? 0), ); ... eventStore.readFromSequence(cleanupStart, Number.MAX_SAFE_INTEGER),
My nine existing projectors were all parked at 147446 (the max sequence) and had been fine for weeks. The new projector's
0won theMath.min, so the upgrade replayed my entire log from sequence 1 and died on page 2.That makes this a landmine for any install carrying a single odd row anywhere in its history, no matter how old or how long it has been harmlessly skipped. The row that broke me was written 17 days before the upgrade and never mattered until v0.0.42 shipped a reader that starts at zero.
2. The undecodable field is in
metadata, nottype.Repro A above poisons
event_type. Mine is a nested optional four levels down:at [25]["metadata"]["origin"]["surface"]origin.surfacewas"migration".ClientSurface(packages/contracts/src/baseSchemas.ts:133) has never accepted that value — v0.0.37 was["web", "desktop", "mobile"], v0.0.42 andmainare["web", "desktop", "mobile", "cli"]. So suggested fix #1 (filter onapplication_event_version) would not have saved me: these are ordinary v1 rows with valid v1 event types. Only suggested fix #2, skip-or-quarantine, covers this — and it needs to catch decode failures anywhere in the row, including insidemetadata, not just unknowntypevalues.Worth noting
mainis unchanged on both counts as of today, and v0.0.42 is current stable, so there is no version to upgrade into.Steps to reproduce
On a v0.0.42 install whose projectors are all caught up, insert one row with an out-of-union nested metadata value, well behind the head:
UPDATE orchestration_events SET metadata_json = json_set(metadata_json, '$.origin.surface', 'migration') WHERE sequence = 1045; DELETE FROM projection_state WHERE projector = 'projection.attachment-cleanup';
Start the app. The new projector bootstraps at
0, replays from the start, and exits 1 on the page containing that row. Every subsequent boot does the same, because the watermark cannot advance. Deleting theprojection_staterow is only there to simulate the upgrade; on a real 0.0.37 → 0.0.42 upgrade it is created at0for you, which is the whole problem.How the bad rows got there
An import of Codex/VS Code session history on 2026-08-31, run under v0.0.37, which stamped 549 events with
origin.surface: "migration"andsource: "codex-vscode":{"origin":{"surface":"migration","appVersion":"0.0.37"},"source":"codex-vscode"}To be clear about provenance, because it affects how much of this is yours: these rows were written directly into SQLite, not through T3's write path. Two independent tells:
AppendEventRequestSchemaencodesmetadataJsonthroughOrchestrationEventMetadata, sosurface: "migration"could never have survivedappend()— it would have failed at encode time.- Their
command_idismigration:codex-vscode:…, which starts with neitherprovider:norserver:, soinferActorKindwould have stamped themclient. They are stored asserver.
So T3 never accepted this write, and
codex-vscodeappears nowhere in this repo. The import was run by an agent against the database file directly. I am not asking you to own the bad data.What I am asking you to own is that one out-of-band row, from any source, now permanently bricks the app on a stock upgrade with no path out short of hand-editing SQLite — and that the error names neither the sequence number nor the table, so finding it took a source clone and a page-arithmetic calculation. An append-only log that is readable at boot only if every historical row is perfectly well-formed has no tolerance budget at all.
Evidence
# Crash loop, 52 occurrences, ~40s apart, from the desktop backend child backend child process failure output start pid=119968 port=3773 backend child process output stdout [13:58:43.263] ERROR (#3): PersistenceDecodeError: ... backend child process failure output end ... repeats; desktop logged "backend exited unexpectedly; restart scheduled" x16 ... # Meanwhile the desktop polls a port that never opens, ~110ms apart, forever GET http://127.0.0.1:3773/.well-known/t3/environment # Projector watermarks at the time of failure: the new one is the odd one out projector last_applied_sequence updated_at projection.projects 147446 2026-09-17T18:43:10.073Z projection.threads 147446 2026-09-17T18:43:10.073Z (... 7 more, all 147446 ...) projection.attachment-cleanup 0 1970-01-01T00:00:00.000Z # The poison: 549 rows, all from the 2026-08-31 import, all valid v1 event types surface n min_seq max_seq (absent) 143905 7 147446 desktop 2327 1 147444 migration 549 1045 1692 <- rejected web 146 13804 145907 event_type n min_seq max_seq occurred_at (first) thread.created 277 1045 1692 2026-04-21T18:52:57.443Z thread.settled 272 1046 1690 2026-04-21T19:08:09.120Z # Page arithmetic, READ_PAGE_SIZE = 500, replay starting at 0 row 526 = sequence 1045 = index 25 of page 2 -> matches "at [25]" exactly # Not corruption PRAGMA quick_check -> ok freelist_count -> 0 orchestration_events -> 146,927 rows, ~1.5 GB of payload+metadata, 3.3 GB databaseFix that worked
Dropping the invalid optional key (rather than substituting a wrong-but-accepted value) was enough:
UPDATE orchestration_events SET metadata_json = json_remove(metadata_json, '$.origin.surface') WHERE json_extract(metadata_json, '$.origin.surface') = 'migration';
549 rows changed. The server came up on the next launch, and
projection.attachment-cleanupcompleted its full 146,927-event replay and advanced 0 → 147446.One note relevant to #10872 / #11182: that full replay is unconditional on first upgrade, and on this database it streams ~1.5 GB of payloads. It survived here, but every v0.0.42 upgrade with a large log is taking that same unbounded
Number.MAX_SAFE_INTEGERread whether or not it has anything to clean up.Environment
- T3 Code 0.0.42 (current stable), upgraded in-place from the desktop app's update prompt
- Windows 11 Pro 26100, x64, Node v26.8.2, sqlite3 3.44.4
- Desktop app against a local server,
serverExposureMode: network-accessible - Providers: codex, claudeAgent, opencode
- No prior fork or nightly build on this machine; migration ledger is unmodified and matches stock
Related
- [Bug]: T3 writes origin.surface="cli" that its own event-store decoder rejects, making t3code unbootable #8789 — same shape,
origin.surfaceagain, fixed by widening the union rather than by making the read survivable. That is why this recurred with a different value. - [Bug]: Nightly 1400/1414 backend OOMs on startup replaying full event log for projection.attachment-cleanup #10872, [Bug]: Startup OOM with oversized replay batches remains unaddressed by #10777 #11182 — the same new projector's full-log replay.
Filed via
t3 triageby Claude Opus 5 (1M context) running as the Claude Code triage agent.Another recurrence on macOS arm64, desktop v0.0.42 (stable, notarized DMG) — no fork builds on my side that I know of.
Symptom: after installing, the app never opens a usable window. The backend child crash-loops about every 30s with exit code 1:
PersistenceSqlError: SQL error in ProjectionSnapshotQuery.getCommandReadModel:listThreads:query: SQLITE(1) SQL logic error [cause]: Error: no such column: pin_order_keyThe desktop app keeps polling
/.well-known/t3/environmentand gets a transport error each time (see also #10517: nothing tells the user what's wrong).Ledger vs code: the DB was created 2026-07-31. Some other build wrote IDs 35–38 on 2026-08-29 under different names. Then v0.0.42 applied 39–52 on 2026-09-28 and skipped 35–38, because only the highest ID is compared:
ID name in DB (written) name in 0.0.42 34 ProjectionThreadsSnoozed (07-31) ProjectionThreadsSnoozed 35 ProjectionThreadsSnoozed (08-29) ProjectionThreadTitleRegeneration 36 ProjectionProjectsCustomTitle (08-29) ProjectionThreadsPinned 37 ProjectionProjectsDefaultThreadEnvMode (08-29) ProjectionTurnsKeysetIndex 38 ProjectionThreadsPinned (08-29) ProjectionThreadsPinOrderKey Note that 35 and 39 are recorded with names from other IDs too (34 and 37 are duplicated). So
038_ProjectionThreadsPinOrderKeynever ran, even though it's idempotent (it checksPRAGMA table_infofirst). Re-running migrations whose recorded name doesn't match the bundled one would have fixed this on its own.Workaround: quit the app, rename
~/.t3/userdata/state.sqlite*aside, relaunch. My DB had no projects, threads or events, so nothing was lost. Users with real history can't use this workaround, though.
What happened
Running either
t3 startort3 servefailed immediately with:The desktop app crash-looped its backend child every ~40 seconds, exit code 1 each time, and never came up.
The trigger was mine: I built an AppImage from a community fork branch to try Copilot CLI support, not realising that branch was based on the unreleased
orchestration-v2work. But the fork only exposed the problem. Once that build had touched the database, stable v0.0.37 was permanently unbootable and could not repair itself, and the two mechanisms that let that happen are both in this repo and both fail silently.Diagnosis
Both failures were reproduced from scratch on a clean, fully migrated v0.0.37 install (details under Steps to reproduce), so neither depends on the fork.
1.
readFromSequencedecodes rows it never filters, and one bad row is terminalapps/server/src/persistence/Layers/OrchestrationEventStore.ts:185-188reads the event log with no discriminator beyond the sequence cursor:The whole page is then decoded as a batch against
OrchestrationEventPersistedRowSchema, whosetypeisOrchestrationEventType. Any row whoseevent_typeis outside that union fails the entire batch.Because
ProjectionPipeline.ts:1778-1779starts each projector from its storedlastAppliedSequence, and that watermark only advances when a batch decodes successfully, the same poisoned batch is re-read on every boot forever. There is no skip, no quarantine, and no way out without hand-editing SQLite.In my case the projectors were parked at 4626 and the very next page was:
event_typethread.createdmessage.updatedwhich is exactly the reported
at [1]["type"].2.
Migratormatches applied migrations by ID and never by nameMigrator.runtakes only the highest applied ID and skips everything at or below it (effect-smolpackages/effect/src/unstable/sql/Migrator.ts:250):Names are recorded in
effect_sql_migrationsbut never compared. So any build that occupies your ID range marks your own migrations as done. My ledger held:OrchestrationV2AuthSessionClientConnectionOrchestrationV2SubagentsProjectionThreadLinkedPullRequestOrchestrationV2FoundationProjectionThreadsUnsettledAtOrchestrationV2*,ApplicationEventSource,ScheduledTasks,LegacyV1ImportStateIDs 1-40 matched v0.0.37 byte for byte, so the divergence began exactly where the two lines forked. v0.0.37's own 041/042/043 therefore never ran, and these columns were simply absent:
auth_sessions.client_surface,auth_sessions.client_app_versionprojection_threads.linked_pull_request_json,projection_threads.unsettled_atNothing reported this. The
Migrations ran successfullyline is only logged when migrations actually execute, so a fully skipped set produces no log output at all. Once the decode crash was cleared, the very first thread-list query died on it:So this is a second, independent brick behind the first.
Why the two combine badly
ApplicationEventSourceaddsorchestration_events.application_event_versionplus the indexidx_orchestration_events_application_sequenceon(application_event_version, sequence)— an index that exists precisely so readers can filter by that column. The v1 reader never learned to use it, so the two event generations share one table with nothing separating them.Worth noting for anyone assuming version numbers are a safety net: the build that did this self-reported as 0.0.33, i.e. older than the 0.0.37 I moved back to. A version string says nothing about which migration line a database is on.
Steps to reproduce
Neither repro needs the fork. Both were run against a clean v0.0.37 install created in a scratch
HOME, and both fail on every subsequent boot, not just the first.A. One unknown event type bricks the server (reproduces the reported error exactly)
On an otherwise empty, fully migrated database, insert a single row:
Then
t3 serve. Result: exit 1 with the identicalreadFromSequence:decodeRowserror, and noserver is ready. Repeated twice — same outcome, since the watermark cannot advance. One row out of one is enough; so is one row out of 4,624.B. Three phantom ledger rows silently disable three migrations
Rewind a clean database to the pre-041 state and let a foreign line claim those IDs:
Then
t3 serve. Result: exit 1 onno such column: linked_pull_request_json, with no migration-related log line whatsoever. All four columns are still missing afterwards. The migrator considers the schema current.How I actually got there (context, not required to reproduce)
T3-Code-0.0.33-x86_64.AppImagefromcopilot-v2on a public fork, to try Copilot CLI support. That branch carries migrations041_OrchestrationV2…049_LegacyV1ImportState— the exact nine names, at the exact nine IDs, now in my ledger. Its PR against this repo has since been closed as not ready.LegacyV1ImportStateappended 402 rows toorchestration_eventstaggedapplication_event_version = 2, using v2-only types:message.updated×197,turn-item.updated×197,thread.metadata-updated×2,thread.visited×2, plus 4 rows whose type names happen to be valid in v1.On the same
orchestration-v2branch today these migrations are renumbered 044-052, i.e. rebased past main's 041-043. The collision was a snapshot-in-time hazard, which is exactly why detection matters rather than discipline.Suggested fix
readFromSequence(andreadAll) anapplication_event_versionpredicate;idx_orchestration_events_application_sequencealready exists for it.no such columncrash six steps downstream. If upstreamMigratorwill not do it, a cheap preflight againstmigrationManifestinMigrations.tswould.Migrations ran successfullyonly appears when something ran; silence currently means both "nothing to do" and "everything was skipped".Happy to split this into separate issues for the reader and the migrator if you would rather track them apart — I filed together because one database ends up wedged by both at once, and either alone would have been recoverable.
Version
0.0.37 (crash). Database previously migrated by a self-built AppImage labelled 0.0.33 from a fork's
copilot-v2branch.Environment
Linux x64, kernel 7.0.0-30-generic, Node 24.20.0, sqlite3 3.46.1. Desktop AppImage against a local server, plus
t3 servefrom a terminal; both fail identically. No systemd service. Providers: claude, copilot.Evidence
Related issues
#8789 — same function and same "startup decoder rejects a persisted row" shape, but triggered by an out-of-enum
origin.surfacevalue rather than an out-of-unionevent_type, and it does not involve the migrator. Not a duplicate; its suggested fix 3 would have prevented this crash, which is why I would treat both as one class. #4518 is the same class again for a settings value. #4374 covers nightly-to-stable downgrade trouble but reports hidden chats, not a crash. #7537 (closed) also concernsreadFromSequencelimits, but the capping behaviour, not decoding.Fix applied or workaround
Rolled the foreign migration line back out of the database, after an integrity-checked
sqlite3 .backupsnapshot, in one transaction:Reading the nine branch migrations first was necessary rather than optional:
ApplicationEventSourcealso addscommand_typetoorchestration_command_receipts, an existing v1 table, which is easy to miss. Reassuringly it only everINSERTs intoorchestration_events— neverUPDATEorDELETE— so no original event row was modified and the v1 log came back intact.v0.0.37 then applied
41_AuthSessionClientConnection,42_ProjectionThreadLinkedPullRequest,43_ProjectionThreadsUnsettledAtand started cleanly; verified over two consecutive cold boots with zero decode errors. All 4,624 v1 events, 2 projects, 2 threads and 197 messages survived.This needed a schema-level diff of the two migration lines to get right. A user without that would reasonably conclude the profile was lost.
Filed by
Claude Fable 5 (
claude-fable-5) viat3 triage, running in Claude Code.