Skip to content

[Bug]: background service self-update fails permanently against a pre-#11510 service-launcher.mjs #11934

Description

@quinnypig

Area

apps/server

Summary

In-app self-update never rewrites ~/.t3/runtime/service-launcher.mjs, so every host running the background service is still using whichever launcher its last CLI-driven install wrote. #11510 changed how the launcher resolves a runtime's entry point, and #11607 changed what the installed runtime tree contains — together they make those persisted launchers reject every new release.

The compatibility gate for exactly this case already exists (runServicePreflight blocks on a launcherProtocol mismatch with "This release requires a newer T3 Code service launcher"), but SERVICE_LAUNCHER_PROTOCOL was left at 2 through #11510, so it never fires. The user gets a misleading error instead.

Steps to reproduce

  1. Install the background service on Linux on a build predating feat(server): manage runtimes as release archives only, never from npm #11510 (my host's service-launcher.mjs was written 2026-08-10) and let it self-update through the app for a few weeks. The launcher file is never replaced.
  2. Let the server reach 0.0.41-nightly.20260914.1722 — the last release whose runtime archive still contained node_modules/t3/dist/bin.mjs.
  3. From any client, run Update server to 0.0.41-nightly.20260915.1735 or newer.

Expected behavior

Either the update lands, or it is blocked up front with the actionable message that servicePreflight.ts already contains: "This release requires a newer T3 Code service launcher. Update it on the server machine."

Actual behavior

The update fails, identically and permanently, on every attempt:

  • The new runtime downloads and installs completely, sentinel and all.
  • The server executes the staged binary's preflight, which returns ready.
  • The launcher then rejects the handoff, and the app shows: Server update failed: The requested target runtime is missing or incomplete.

That message is wrong in a way that costs debugging time: the target runtime is complete and was just successfully executed by the same code path. What is actually stale is the launcher.

The mechanism:

0.0.41-nightly.20260914.1722 shipped both layouts, which is why it was the last version to install successfully.

Impact

Major degradation or frequent failure

Silent and self-perpetuating: the update looks like it is working (the version installs fully), the error suggests a corrupt download, and nothing points at the real fix. It should hit any background-service host whose launcher predates #11510, which is effectively all of them, since in-app self-update never replaces that file.

Version or commit

Server stuck on 0.0.41-nightly.20260914.1722; failing to reach 0.0.41-nightly.20260915.1735 and .1766. Analysis against main @ 7235701.

Environment

Ubuntu 24.04 arm64, systemd user service, nightly channel, headless (no desktop app).

Logs or stack traces

# ~/.t3/userdata/logs/server.trace.ndjson, span cloud.server_self_update.update
ServerSelfUpdateError: Server update failed: The requested target runtime is missing or incomplete.
    at failWith (.../versions/0.0.41-nightly.20260914.1722/t3:87571:97)
    at cloud.server_self_update.update (.../versions/0.0.41-nightly.20260914.1722/t3:87533:40) {
  [cause]: ServiceLauncherRejectedError: The requested target runtime is missing or incomplete.
}

# state after two failed attempts: new runtimes fully installed, never activated
$ ls ~/.t3/runtime/versions/
0.0.41-nightly.20260914.1722   <- active
0.0.41-nightly.20260915.1735   <- installed, valid .install-complete, never used
0.0.41-nightly.20260915.1766   <- installed, valid .install-complete, never used

$ cat ~/.t3/runtime/service-state.json
{"protocol":2,"activeVersion":"0.0.41-nightly.20260914.1722","update":{...,"status":"committed"}}
# untouched since the .1722 update; the launcher never records a pending update

$ ls -l ~/.t3/runtime/service-launcher.mjs
-rw-rw-r-- 1 ... 21382 Aug 10 17:39   # never rewritten by ~25 successful self-updates
$ grep -n 'entryPath' ~/.t3/runtime/service-launcher.mjs
116: entryPath: NodePath.join(versionDir, "node_modules", "t3", "dist", "bin.mjs"),

Note: ~/.t3/userdata/logs/boot-service.log had been deleted on disk while systemd still held the fd, so the launcher-side log was only reachable via /proc/<launcher-pid>/fd/1. Possibly worth a separate look — nothing in the repo appears to unlink it.

Workaround

Run t3 update on the server machine from a current on-disk binary. This rewrites the unit to ExecStart=<versionDir>/t3 __service-launcher, which hosts the launcher inside the versioned executable and fixes it permanently:

~/.t3/runtime/versions/<newest>/t3 update --channel nightly --yes

Confirmed working — the service came back on .1766, with drop-ins and Tailscale Serve intact.

Suggested fix

Bump SERVICE_LAUNCHER_PROTOCOL to 3. The preflight then blocks before anything downloads and emits the correct message. Worth considering alongside it: any layout change to runtimePaths is a launcher compatibility break by definition, since that file outlives every self-update, so a comment on the constant tying the two together would help the next person.

Activity

  1. juliusmarminge commented on Sep 15, 2026

    @juliusmarminge
    Member

    Triage

    Confirmed on main @ 7235701de0. This is a real bug, not a duplicate of #8345 / #11767 / #5196 / #9370.

    In-app Update server never rewrites the boot-service unit or ~/.t3/runtime/service-launcher.mjs. That is intentional: docs/internals/server-updates.md says children request a handoff over IPC and do not replace their service definition. A layout change therefore has to trip SERVICE_LAUNCHER_PROTOCOL, or an old launcher will keep supervising new runtimes with a stale entryPath.

    What the code does today

    .1722 was the last version that installed because #11732 wrote a legacy node_modules/t3/dist/bin.mjs into archives. #11770 removed that the same day, on the assumption that “an old launcher never installs an archive.” That is wrong for this path: the running server downloads the archive; the persisted launcher only has to accept the handoff. .1735 / .1766 therefore look complete to the installer and incomplete to the old launcher.

    The npm-only shim in legacyCliLauncher.ts does not help archive self-update. Its comment that “the first server started this way rewrites the service unit” is also false for the in-app path.

    Workaround (confirmed by reporter)

    On the server machine, from a current on-disk binary:

    ~/.t3/runtime/versions/<newest>/t3 update --channel nightly --yes

    That rewrites the unit to ExecStart=<versionDir>/t3 __service-launcher and unblocks later in-app updates.

    Suggested fix

    1. Bump SERVICE_LAUNCHER_PROTOCOL to 3 and comment that any runtimePaths / installed-tree change is a launcher compatibility break. The next nightly’s preflight will then block with the existing actionable message instead of the false “runtime missing” error. Users still need a local t3 update to actually move — that is the designed escape hatch.
    2. Do not treat restoring fix(release): preserve updates from npm-based services #11732’s archive ensureLegacyEntry as the durable fix. It would let old .mjs launchers succeed once more, but it leaves them in place and fights chore(server): keep the legacy service entry point to the npm package only #11770.

    boot-service.log vanishing while systemd still holds the fd looks separate: nothing in-repo unlinks that file.

    No open PR covers this. Accepting as a bug.

  2. added
    bugSomething is broken or behaving incorrectly.
    acceptedfeature request accepted
    via-triageFiled through npx t3 triage
    on Sep 15, 2026
  3. iniTwakkie commented on Sep 16, 2026

    @iniTwakkie

    Independent confirmation on the stable 0.0.42 release, with an additional systemd drop-in trigger that may not be covered by #11940.

    Environment:

    Ubuntu 24.04.5 LTS x86_64
    kernel 6.8.0-139-generic
    Node v24.21.0
    systemd user service
    command: npx --yes t3@latest service update --base-dir ~/.t3
    

    Observed sequence:

    1. service update installed the complete standalone 0.0.42 runtime and rewrote the base unit to:

      ExecStart=~/.t3/runtime/versions/0.0.42/t3 __service-launcher
    2. An existing local drop-in still contained:

      ExecStart=
      ExecStart=%h/.local/bin/node %h/.t3/runtime/service-launcher.mjs
    3. Consequently, systemctl --user show t3code.service -p ExecStart still resolved to the legacy Node launcher after the update. service-state.json selected 0.0.42, whose valid standalone layout intentionally has no node_modules/t3/dist/bin.mjs.

    4. The legacy launcher rejected the active runtime as missing/incomplete. systemd attempted five starts and left the service failed with no listener. The update command nevertheless logged completion.

    5. Rolling activeVersion back to 0.0.40, resetting the failed unit, and restarting restored the service immediately.

    This confirms the compatibility failure against stable 0.0.42. It also leaves a question beyond the in-app preflight fixed by #11940: does a local service update validate the effective systemd ExecStart after daemon-reload? In this case the base unit was correct, but an ExecStart= drop-in silently defeated it and the CLI still reported success.

    A possible hardening would be for service install/update to inspect the effective ExecStart after writing/reloading the unit and either refuse activation or report the overriding drop-in path when the expected standalone launcher is not effective.

  4. vtsixthai commented on Sep 16, 2026

    @vtsixthai

    Hit this too, from a slightly different angle. I run a small script on a cron that connects to the server over a WebSocket, and it was borrowing ws straight out of ~/.t3/runtime/versions/<newest>/node_modules. After the auto-update to 0.0.42 that folder no longer has ws, so the script died with Cannot find module .../0.0.42/node_modules/ws.

    Makes sense in hindsight now that the server ships as a single binary with ws bundled inside. I've fixed my side by installing ws as a normal dependency of my own script instead of reaching into the runtime folder. No action needed from me, just mentioning it in case other folks have external scripts that resolve packages from the runtime directory. Might be worth a line in the release notes if it isn't already.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    acceptedfeature request acceptedbugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions