Skip to content

Colliding migration IDs silently skip migrations; one undecodable event row bricks startup #8896

Description

@hackel

What happened

Running either t3 start or t3 serve failed immediately with:

ERROR (#5): PersistenceDecodeError: Decode error in OrchestrationEventStore.readFromSequence:decodeRows: Composite(Pointer(Composite(Pointer(AnyOf()))))

The desktop app crash-looped its backend child every ~40 seconds, exit code 1 each time, and never came up.

The trigger was mine: I built an AppImage from a community fork branch to try Copilot CLI support, not realising that branch was based on the unreleased orchestration-v2 work. But the fork only exposed the problem. Once that build had touched the database, stable v0.0.37 was permanently unbootable and could not repair itself, and the two mechanisms that let that happen are both in this repo and both fail silently.

Diagnosis

Both failures were reproduced from scratch on a clean, fully migrated v0.0.37 install (details under Steps to reproduce), so neither depends on the fork.

1. readFromSequence decodes rows it never filters, and one bad row is terminal

apps/server/src/persistence/Layers/OrchestrationEventStore.ts:185-188 reads the event log with no discriminator beyond the sequence cursor:

FROM orchestration_events
WHERE sequence > ${request.sequenceExclusive}
ORDER BY sequence ASC
LIMIT ${request.limit}

The whole page is then decoded as a batch against OrchestrationEventPersistedRowSchema, whose type is OrchestrationEventType. Any row whose event_type is outside that union fails the entire batch.

Because ProjectionPipeline.ts:1778-1779 starts each projector from its stored lastAppliedSequence, and that watermark only advances when a batch decodes successfully, the same poisoned batch is re-read on every boot forever. There is no skip, no quarantine, and no way out without hand-editing SQLite.

In my case the projectors were parked at 4626 and the very next page was:

index sequence event_type in the v1 union?
0 4627 thread.created yes
1 4628 message.updated no

which is exactly the reported at [1]["type"].

2. Migrator matches applied migrations by ID and never by name

Migrator.run takes only the highest applied ID and skips everything at or below it (effect-smol packages/effect/src/unstable/sql/Migrator.ts:250):

if (currentId <= latestMigrationId) {
  continue
}

Names are recorded in effect_sql_migrations but never compared. So any build that occupies your ID range marks your own migrations as done. My ledger held:

ID name in my DB name in v0.0.37
41 OrchestrationV2 AuthSessionClientConnection
42 OrchestrationV2Subagents ProjectionThreadLinkedPullRequest
43 OrchestrationV2Foundation ProjectionThreadsUnsettledAt
44-49 OrchestrationV2*, ApplicationEventSource, ScheduledTasks, LegacyV1ImportState (do not exist)

IDs 1-40 matched v0.0.37 byte for byte, so the divergence began exactly where the two lines forked. v0.0.37's own 041/042/043 therefore never ran, and these columns were simply absent:

  • auth_sessions.client_surface, auth_sessions.client_app_version
  • projection_threads.linked_pull_request_json, projection_threads.unsettled_at

Nothing reported this. The Migrations ran successfully line is only logged when migrations actually execute, so a fully skipped set produces no log output at all. Once the decode crash was cleared, the very first thread-list query died on it:

PersistenceSqlError: SQL error in ProjectionSnapshotQuery.getCommandReadModel:listThreads:query
  [cause]: Error: no such column: linked_pull_request_json

So this is a second, independent brick behind the first.

Why the two combine badly

ApplicationEventSource adds orchestration_events.application_event_version plus the index idx_orchestration_events_application_sequence on (application_event_version, sequence) — an index that exists precisely so readers can filter by that column. The v1 reader never learned to use it, so the two event generations share one table with nothing separating them.

Worth noting for anyone assuming version numbers are a safety net: the build that did this self-reported as 0.0.33, i.e. older than the 0.0.37 I moved back to. A version string says nothing about which migration line a database is on.

Steps to reproduce

Neither repro needs the fork. Both were run against a clean v0.0.37 install created in a scratch HOME, and both fail on every subsequent boot, not just the first.

A. One unknown event type bricks the server (reproduces the reported error exactly)

On an otherwise empty, fully migrated database, insert a single row:

INSERT INTO orchestration_events
 (event_id, aggregate_kind, stream_id, stream_version, event_type, occurred_at,
  command_id, causation_event_id, correlation_id, actor_kind, payload_json, metadata_json)
VALUES
 ('repro-unknown-type','thread','00000000-0000-4000-8000-000000000001',0,
  'message.updated','2026-08-31T00:00:00.000Z',NULL,NULL,NULL,'server','{}','{}');

Then t3 serve. Result: exit 1 with the identical readFromSequence:decodeRows error, and no server is ready. Repeated twice — same outcome, since the watermark cannot advance. One row out of one is enough; so is one row out of 4,624.

B. Three phantom ledger rows silently disable three migrations

Rewind a clean database to the pre-041 state and let a foreign line claim those IDs:

DELETE FROM effect_sql_migrations WHERE migration_id >= 41;
ALTER TABLE projection_threads DROP COLUMN unsettled_at;
ALTER TABLE projection_threads DROP COLUMN linked_pull_request_json;
ALTER TABLE auth_sessions DROP COLUMN client_surface;
ALTER TABLE auth_sessions DROP COLUMN client_app_version;
INSERT INTO effect_sql_migrations (migration_id, created_at, name) VALUES
 (41,'2026-08-28 13:45:59','OrchestrationV2'),
 (42,'2026-08-28 13:45:59','OrchestrationV2Subagents'),
 (43,'2026-08-28 13:45:59','OrchestrationV2Foundation');

Then t3 serve. Result: exit 1 on no such column: linked_pull_request_json, with no migration-related log line whatsoever. All four columns are still missing afterwards. The migrator considers the schema current.

How I actually got there (context, not required to reproduce)

  1. Built T3-Code-0.0.33-x86_64.AppImage from copilot-v2 on a public fork, to try Copilot CLI support. That branch carries migrations 041_OrchestrationV2 … 049_LegacyV1ImportState — the exact nine names, at the exact nine IDs, now in my ledger. Its PR against this repo has since been closed as not ready.
  2. Ran it once, 2026-08-28 13:45:59. It applied 35-40 (matching main) then 41-49 (its own), and LegacyV1ImportState appended 402 rows to orchestration_events tagged application_event_version = 2, using v2-only types: message.updated ×197, turn-item.updated ×197, thread.metadata-updated ×2, thread.visited ×2, plus 4 rows whose type names happen to be valid in v1.
  3. Went back to stable v0.0.37 — permanently unbootable from then on.

On the same orchestration-v2 branch today these migrations are renumbered 044-052, i.e. rebased past main's 041-043. The collision was a snapshot-in-time hazard, which is exactly why detection matters rather than discipline.

Suggested fix

  1. Filter the read. Give readFromSequence (and readAll) an application_event_version predicate; idx_orchestration_events_application_sequence already exists for it.
  2. Make an undecodable row survivable. Skip or quarantine it with a loud warning instead of failing the batch. As [Bug]: T3 writes origin.surface="cli" that its own event-store decoder rejects, making t3code unbootable #8789 argues, a single malformed event in an append-only log should never make the service unbootable — and because the watermark is gated on batch success, today it is unbootable permanently.
  3. Verify migration identity, not just height. Compare recorded names against loaded ones for every applied ID and refuse to start on a mismatch, naming the offending IDs. A clear "this database was migrated by a different build" beats a no such column crash six steps downstream. If upstream Migrator will not do it, a cheap preflight against migrationManifest in Migrations.ts would.
  4. Log the no-op case. Migrations ran successfully only appears when something ran; silence currently means both "nothing to do" and "everything was skipped".

Happy to split this into separate issues for the reader and the migrator if you would rather track them apart — I filed together because one database ends up wedged by both at once, and either alone would have been recoverable.

Version

0.0.37 (crash). Database previously migrated by a self-built AppImage labelled 0.0.33 from a fork's copilot-v2 branch.

Environment

Linux x64, kernel 7.0.0-30-generic, Node 24.20.0, sqlite3 3.46.1. Desktop AppImage against a local server, plus t3 serve from a terminal; both fail identically. No systemd service. Providers: claude, copilot.

Evidence

# Reported failure, every boot (paths and home directory redacted)
[08:25:01.241] ERROR (#5): PersistenceDecodeError: Decode error in
  OrchestrationEventStore.readFromSequence:decodeRows: Composite(Pointer(Composite(Pointer(AnyOf()))))
  [cause]: SchemaError: Expected "project.created" | "project.meta-updated" | ... | "thread.activity-appended"
    at [1]["type"]

# Desktop backend crash-loop, ~40s apart, all exit 1
12:55:32 backend child process failure output start  pid=... port=3773
12:55:32 backend child process output  stdout  [07:55:03.439] ERROR (#5): PersistenceDecodeError: ...
12:55:32 backend child process failure output end    code=1
... repeats through 13:04:52 ...

# Event types present, by application_event_version. v=2 rows were unreadable by v0.0.37.
v  event_type                            n     min_seq  max_seq
1  thread.activity-appended              3285  19       4623
1  thread.message-sent                   840   6        4618
1  thread.session-set                    468   8        4624
1  (9 further v1 types)                  ...
2  project.created                       2     4625     4626     <- valid name, from the import
2  thread.created                        2     4627     4633     <- valid name, from the import
2  message.updated                       197   4628     5023     <- rejected
2  turn-item.updated                     197   4629     5024     <- rejected
2  thread.metadata-updated               2     4632     4638     <- rejected (v1 has thread.meta-updated)
2  thread.visited                        2     5025     5026     <- rejected

# All nine projectors parked immediately before the first rejected row
projector                         last_applied_sequence
projection.projects               4626
projection.threads                4626
(... 7 more, all 4626)

# Ledger: IDs 1-40 match v0.0.37 exactly; 41-49 are a different migration line
41  2026-08-28 13:45:59  OrchestrationV2                          <- v0.0.37 expects AuthSessionClientConnection
42  2026-08-28 13:45:59  OrchestrationV2Subagents                 <- v0.0.37 expects ProjectionThreadLinkedPullRequest
43  2026-08-28 13:45:59  OrchestrationV2Foundation                <- v0.0.37 expects ProjectionThreadsUnsettledAt
44  2026-08-28 13:45:59  OrchestrationV2ProviderSessionBindings
45  2026-08-28 13:45:59  OrchestrationV2ThreadLaunchWorkflows
46  2026-08-28 13:45:59  ApplicationEventSource
47  2026-08-28 13:45:59  OrchestrationV2EffectCancellation
48  2026-08-28 13:45:59  ScheduledTasks
49  2026-08-28 13:45:59  LegacyV1ImportState

# Second brick, surfaced only after the decode crash was cleared
ERROR (#5): PersistenceSqlError: SQL error in
  ProjectionSnapshotQuery.getCommandReadModel:listThreads:query
  [cause]: Error: no such column: linked_pull_request_json

# Proof the three migrations really had not run
projection_threads.unsettled_at            : absent
projection_threads.linked_pull_request_json: absent
auth_sessions.client_surface               : absent
auth_sessions.client_app_version           : absent

Related issues

#8789 — same function and same "startup decoder rejects a persisted row" shape, but triggered by an out-of-enum origin.surface value rather than an out-of-union event_type, and it does not involve the migrator. Not a duplicate; its suggested fix 3 would have prevented this crash, which is why I would treat both as one class. #4518 is the same class again for a settings value. #4374 covers nightly-to-stable downgrade trouble but reports hidden chats, not a crash. #7537 (closed) also concerns readFromSequence limits, but the capping behaviour, not decoding.

Fix applied or workaround

Rolled the foreign migration line back out of the database, after an integrity-checked sqlite3 .backup snapshot, in one transaction:

DELETE FROM orchestration_events WHERE application_event_version = 2;       -- 402 rows
DELETE FROM orchestration_command_receipts WHERE command_type <> 'legacy';  -- 2 rows
UPDATE projection_state SET last_applied_sequence = 4624 WHERE last_applied_sequence > 4624;
DROP INDEX IF EXISTS idx_orchestration_events_application_sequence;
ALTER TABLE orchestration_events DROP COLUMN application_event_version;
ALTER TABLE orchestration_command_receipts DROP COLUMN command_type;
DROP TABLE ...;                          -- 24 orchestration_v2_* tables, plus scheduled_tasks
DELETE FROM effect_sql_migrations WHERE migration_id >= 41;
UPDATE sqlite_sequence SET seq = 4624 WHERE name = 'orchestration_events';

Reading the nine branch migrations first was necessary rather than optional: ApplicationEventSource also adds command_type to orchestration_command_receipts, an existing v1 table, which is easy to miss. Reassuringly it only ever INSERTs into orchestration_events — never UPDATE or DELETE — so no original event row was modified and the v1 log came back intact.

v0.0.37 then applied 41_AuthSessionClientConnection, 42_ProjectionThreadLinkedPullRequest, 43_ProjectionThreadsUnsettledAt and started cleanly; verified over two consecutive cold boots with zero decode errors. All 4,624 v1 events, 2 projects, 2 threads and 197 messages survived.

This needed a schema-level diff of the two migration lines to get right. A user without that would reasonably conclude the profile was lost.

Filed by

Claude Fable 5 (claude-fable-5) via t3 triage, running in Claude Code.

Activity

  1. hackel commented on Aug 31, 2026

    @hackel
    Author

    Possible dupe of #8789.

  2. ImBIOS commented on Sep 3, 2026

    @ImBIOS

    Independent recurrence of mechanism 2 (migrator half), this time on official nightly 0.0.39 with no fork build involved in the skipping step.

    Ledger vs code (production DB, journal timestamps UTC):

    ID name in DB name in 0.0.39
    42 ProjectionThreadLinkedPullRequest ProjectionThreadLinkedPullRequest
    43 ProjectionThreadsUnsettledAt ProjectionThreadsUnsettledAt
    44 AuthSessionClientInstanceId ClearAutomaticProjectModelDefaults
    45 ProjectionTasks ProjectionProjectsAutoPull
    46 CleanOscPollution RepairAutomaticSettlementTimestamps

    IDs 44–46 in the journal came from an earlier fork line (my own dev branch, applied 2026-08-25); upstream later reused the same slots. Max recorded ID was already 46, so 044/045/046 never ran — zero migration log output — and the backend child crash-looped (exit 1, ~35s period) on:

    PersistenceSqlError ... ProjectionSnapshotQuery.getCommandReadModel:listProjects ... [cause]: Error: no such column: auto_pull

    while the desktop window never became ready. I repaired by hand-applying the three skipped migrations (045 is a guarded ALTER, 044/046 are idempotent UPDATEs) after snapshotting the DB; backend has been stable since, health endpoint reporting 0.0.39.

    Two notes for the fix discussion:

    • The skipped migration here is from feat(projects): automatically pull clean default branches #9277 (merged Sep 2), so this failure mode is reachable by any user whose journal hit 46 through another line — not just fork users.
    • A fatal preflight would have bricked this database outright (44–46 legitimately differ). Repair-by-new-ID plus a loud non-fatal warning seems the safer shape.

    PR with both: #9312 (repair migration 047 + slot-collision warning in runMigrations, with a regression test that encodes exactly this journal state).

  3. joshuadoge commented on Sep 17, 2026

    @joshuadoge

    Hit the same defect on Windows, and I think this case closes the "but you ran a fork build" escape hatch: no fork was involved, and the trigger was a stock upgrade to current stable v0.0.42.

    Reporting via t3 triage.

    What happened

    Desktop app offered an update, I took it, clicked the restart button, and the ADE never came back. Reinstalling and rebooting the machine changed nothing. The backend crash-looped every ~40s with the error from this issue:

    ERROR (#3): PersistenceDecodeError: Decode error in
      OrchestrationEventStore.readFromSequence:decodeRows:
      Composite(Pointer(Composite(Pointer(Composite(Pointer(AnyOf(Composite(Pointer(AnyOf(AnyOf()))))))))))
      [cause]: SchemaError: Expected "web" | "desktop" | "mobile" | "cli"
        at [25]["metadata"]["origin"]["surface"]
    

    Two things this adds to the report above

    1. v0.0.42 turns any pre-existing undecodable row into a boot failure, on the supported upgrade path.

    v0.0.42 adds a projector that does not exist in v0.0.37: projection.attachment-cleanup. It arrives with a fresh watermark of 0, and ProjectionPipeline.ts:2124 then drags the read back to the very beginning regardless of how far everything else has advanced:

    const cleanupStart = Math.min(
      cleanupState?.lastAppliedSequence ?? 0,
      ...projectors.map((projector) => byProjector.get(projector.name)?.lastAppliedSequence ?? 0),
    );
    ...
    eventStore.readFromSequence(cleanupStart, Number.MAX_SAFE_INTEGER),

    My nine existing projectors were all parked at 147446 (the max sequence) and had been fine for weeks. The new projector's 0 won the Math.min, so the upgrade replayed my entire log from sequence 1 and died on page 2.

    That makes this a landmine for any install carrying a single odd row anywhere in its history, no matter how old or how long it has been harmlessly skipped. The row that broke me was written 17 days before the upgrade and never mattered until v0.0.42 shipped a reader that starts at zero.

    2. The undecodable field is in metadata, not type.

    Repro A above poisons event_type. Mine is a nested optional four levels down:

    at [25]["metadata"]["origin"]["surface"]
    

    origin.surface was "migration". ClientSurface (packages/contracts/src/baseSchemas.ts:133) has never accepted that value — v0.0.37 was ["web", "desktop", "mobile"], v0.0.42 and main are ["web", "desktop", "mobile", "cli"]. So suggested fix #1 (filter on application_event_version) would not have saved me: these are ordinary v1 rows with valid v1 event types. Only suggested fix #2, skip-or-quarantine, covers this — and it needs to catch decode failures anywhere in the row, including inside metadata, not just unknown type values.

    Worth noting main is unchanged on both counts as of today, and v0.0.42 is current stable, so there is no version to upgrade into.

    Steps to reproduce

    On a v0.0.42 install whose projectors are all caught up, insert one row with an out-of-union nested metadata value, well behind the head:

    UPDATE orchestration_events
    SET metadata_json = json_set(metadata_json, '$.origin.surface', 'migration')
    WHERE sequence = 1045;
    
    DELETE FROM projection_state WHERE projector = 'projection.attachment-cleanup';

    Start the app. The new projector bootstraps at 0, replays from the start, and exits 1 on the page containing that row. Every subsequent boot does the same, because the watermark cannot advance. Deleting the projection_state row is only there to simulate the upgrade; on a real 0.0.37 → 0.0.42 upgrade it is created at 0 for you, which is the whole problem.

    How the bad rows got there

    An import of Codex/VS Code session history on 2026-08-31, run under v0.0.37, which stamped 549 events with origin.surface: "migration" and source: "codex-vscode":

    {"origin":{"surface":"migration","appVersion":"0.0.37"},"source":"codex-vscode"}

    To be clear about provenance, because it affects how much of this is yours: these rows were written directly into SQLite, not through T3's write path. Two independent tells:

    • AppendEventRequestSchema encodes metadataJson through OrchestrationEventMetadata, so surface: "migration" could never have survived append() — it would have failed at encode time.
    • Their command_id is migration:codex-vscode:…, which starts with neither provider: nor server:, so inferActorKind would have stamped them client. They are stored as server.

    So T3 never accepted this write, and codex-vscode appears nowhere in this repo. The import was run by an agent against the database file directly. I am not asking you to own the bad data.

    What I am asking you to own is that one out-of-band row, from any source, now permanently bricks the app on a stock upgrade with no path out short of hand-editing SQLite — and that the error names neither the sequence number nor the table, so finding it took a source clone and a page-arithmetic calculation. An append-only log that is readable at boot only if every historical row is perfectly well-formed has no tolerance budget at all.

    Evidence

    # Crash loop, 52 occurrences, ~40s apart, from the desktop backend child
    backend child process failure output start  pid=119968 port=3773
    backend child process output  stdout  [13:58:43.263] ERROR (#3): PersistenceDecodeError: ...
    backend child process failure output end
    ... repeats; desktop logged "backend exited unexpectedly; restart scheduled" x16 ...
    
    # Meanwhile the desktop polls a port that never opens, ~110ms apart, forever
    GET http://127.0.0.1:3773/.well-known/t3/environment
    
    # Projector watermarks at the time of failure: the new one is the odd one out
    projector                         last_applied_sequence  updated_at
    projection.projects               147446                 2026-09-17T18:43:10.073Z
    projection.threads                147446                 2026-09-17T18:43:10.073Z
    (... 7 more, all 147446 ...)
    projection.attachment-cleanup     0                      1970-01-01T00:00:00.000Z
    
    # The poison: 549 rows, all from the 2026-08-31 import, all valid v1 event types
    surface     n       min_seq  max_seq
    (absent)    143905  7        147446
    desktop     2327    1        147444
    migration   549     1045     1692     <- rejected
    web         146     13804    145907
    
    event_type      n    min_seq  max_seq  occurred_at (first)
    thread.created  277  1045     1692     2026-04-21T18:52:57.443Z
    thread.settled  272  1046     1690     2026-04-21T19:08:09.120Z
    
    # Page arithmetic, READ_PAGE_SIZE = 500, replay starting at 0
    row 526 = sequence 1045 = index 25 of page 2  -> matches "at [25]" exactly
    
    # Not corruption
    PRAGMA quick_check    -> ok
    freelist_count        -> 0
    orchestration_events  -> 146,927 rows, ~1.5 GB of payload+metadata, 3.3 GB database
    

    Fix that worked

    Dropping the invalid optional key (rather than substituting a wrong-but-accepted value) was enough:

    UPDATE orchestration_events
    SET metadata_json = json_remove(metadata_json, '$.origin.surface')
    WHERE json_extract(metadata_json, '$.origin.surface') = 'migration';

    549 rows changed. The server came up on the next launch, and projection.attachment-cleanup completed its full 146,927-event replay and advanced 0 → 147446.

    One note relevant to #10872 / #11182: that full replay is unconditional on first upgrade, and on this database it streams ~1.5 GB of payloads. It survived here, but every v0.0.42 upgrade with a large log is taking that same unbounded Number.MAX_SAFE_INTEGER read whether or not it has anything to clean up.

    Environment

    • T3 Code 0.0.42 (current stable), upgraded in-place from the desktop app's update prompt
    • Windows 11 Pro 26100, x64, Node v26.8.2, sqlite3 3.44.4
    • Desktop app against a local server, serverExposureMode: network-accessible
    • Providers: codex, claudeAgent, opencode
    • No prior fork or nightly build on this machine; migration ledger is unmodified and matches stock

    Related


    Filed via t3 triage by Claude Opus 5 (1M context) running as the Claude Code triage agent.

  4. ishanavasthi commented on Sep 28, 2026

    @ishanavasthi

    Another recurrence on macOS arm64, desktop v0.0.42 (stable, notarized DMG) — no fork builds on my side that I know of.

    Symptom: after installing, the app never opens a usable window. The backend child crash-loops about every 30s with exit code 1:

    PersistenceSqlError: SQL error in ProjectionSnapshotQuery.getCommandReadModel:listThreads:query: SQLITE(1) SQL logic error
      [cause]: Error: no such column: pin_order_key
    

    The desktop app keeps polling /.well-known/t3/environment and gets a transport error each time (see also #10517: nothing tells the user what's wrong).

    Ledger vs code: the DB was created 2026-07-31. Some other build wrote IDs 35–38 on 2026-08-29 under different names. Then v0.0.42 applied 39–52 on 2026-09-28 and skipped 35–38, because only the highest ID is compared:

    ID name in DB (written) name in 0.0.42
    34 ProjectionThreadsSnoozed (07-31) ProjectionThreadsSnoozed
    35 ProjectionThreadsSnoozed (08-29) ProjectionThreadTitleRegeneration
    36 ProjectionProjectsCustomTitle (08-29) ProjectionThreadsPinned
    37 ProjectionProjectsDefaultThreadEnvMode (08-29) ProjectionTurnsKeysetIndex
    38 ProjectionThreadsPinned (08-29) ProjectionThreadsPinOrderKey

    Note that 35 and 39 are recorded with names from other IDs too (34 and 37 are duplicated). So 038_ProjectionThreadsPinOrderKey never ran, even though it's idempotent (it checks PRAGMA table_info first). Re-running migrations whose recorded name doesn't match the bundled one would have fixed this on its own.

    Workaround: quit the app, rename ~/.t3/userdata/state.sqlite* aside, relaunch. My DB had no projects, threads or events, so nothing was lost. Users with real history can't use this workaround, though.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions