Repository navigation
core: 60m idle location eviction interrupts a running session and rejects pending questions #51343
Description
Activity
github-actions commented
on Sep 25, 2026 on Sep 25, 2026 – with GitHub ActionsContributorMore actionsHello all,
I am having the same issue here. Yesterday I had opencode v1 + openchamber v1 running well with sessions of 3 to 12 hours (I have a very old server graphics card that take up to 3-4 hours to give the first token). Once I updated to opencode v2.0.16 + openchamber 2.0.1 I got the same issue. After 1 hour I got the message:
Opencode failed to send message with error: Step interruptedI tested several times with some parameters in the model (if possible to get the timeout) but it didn't work.
If the model answer in less than 60 minutes everything works correctly.Regards.
Correction to the initial description: I first attributed this to a LayerMap
idleTimeToLive: "60 minutes"inlocation-services.ts, but that was from thedevtree. On the released 2.0.8,location-services.tsbuilds the LayerMap withidleTimeToLive: Duration.infinity; the 60-minute mechanism isLocationActivity(packages/core/src/location-activity.ts), which renews only on durable session events carrying alocationand, on expiry, interrupts active executions withreason: "inactivity". I have updated the description accordingly.Still happens on 2.0.19, and seen here on 2.0.16 and 2.0.18 too (headless
opencode serve, OpenChamber as the client). Five waits for an answer ended 60.2 to 60.6 minutes after the tool started, each on the same second aslocation services evictedfor the session's directory: three on the built-inquestiontool and two on a plugin's review form (Plannotator'ssubmit_plan). So forms created by plugins are cut off the same way. #50499 would cover all five.Corroborating this on 2.0.18 (latest channel) — same failure, with some new data points. My topology differs from the original report:
opencode serve --servicewith the TUI client on the same host (SQLite session storage), and the TUI process stayed alive with an ESTABLISHED socket for the entire window, so this is not a client disconnect.Verified timeline (2026-10-01, UTC; from session storage + server log):
Time (UTC) Event 10:58:49.998 questiontool part starts running, waiting for user input12:29:52.8–53.0 snapshot spawns + location services evicted12:29:52.859 pending questionpart completes with{"type":"aborted","message":"Tool execution interrupted"}— pending for 5462.9 s (~91 min)12:29:53 session idleevent{"outcome":"interrupted"};time_idle/idle_outcomeset13:46:49 user returns and sends a message; only now does the model see the aborted tool result (77 min after the prompt died) Threshold: ~90 min, not 60 — worth reconciling. My prompt was pending 91.0 min before eviction. If part creation/
rancounts as the last durable event, a 60-minute TTL should have fired ~31 min earlier, so either the TTL changed by 2.0.18 (looks like ~90 min now) or the renewal set is broader than described. For what it's worth, an occurrence in another session the previous day evicted only ~62 min after going pending, so my two data points straddle both 60 and ~90.State of the pending question part in the session storage after eviction (on 2.0.18 it's the tool part, not the assistant step, that records the abort):
{ "type": "tool", "name": "question", "executed": false, "state": { "status": "error", "error": { "type": "aborted", "message": "Tool execution interrupted" }, "time": { "created": 1790852326718, "ran": 1790852329998, "completed": 1790857792859 } } }(
created= 10:58:46.718Z,ran= 10:58:49.998Z,completed= 12:29:52.859Z.) No notification of any kind — the prompt is simply dead when the user returns.Server log — note the location is re-booted 335 ms after eviction:
timestamp=2026-10-01T12:29:53.034Z level=INFO run=… message="location services evicted" directory=/…/workspace workspaceID=undefined role=server timestamp=2026-10-01T12:29:53.369Z level=INFO run=… message="location services booted" directory=/…/workspace workspaceID=undefined durationMs=217 http.span=217 role=serverThe abort timestamp (12:29:52.859) precedes the eviction log line by ~175 ms, consistent with the
awaitSettlement: truesequencing below.Eviction path in the shipped 2.0.18 bundle (reformatted, names shortened) — confirms the mechanism described in the original report is core bundle code, no plugin involved:
const now = currentTimeMillisUnsafe() const expired = Array.from(locations.values()).filter((loc) => loc.expiresAt <= now) if (expired.length === 0) return for (const loc of expired) { const runsInLocation = activeRuns.flatMap((run) => sameLocation(run.location, loc.ref) ? [run] : [], ) yield* each(runsInLocation, (run) => sessions.interrupt(run.id, { reason: "inactivity", awaitSettlement: true }), ) locations.delete(locationKey(loc.ref)) log("location services evicted", { directory: loc.ref.directory, workspaceID: loc.ref.workspaceID, }) }
Workaround: ask questions as plain assistant text instead of the interactive
questiontool — a completed turn survives inactivity and the conversation resumes when the user returns. As in the original report, I found no config, env var, or CLI flag to change the TTL or opt out of the interrupt.Happy to provide the full session-storage dump or additional log context if useful.
Consolidating what the open PRs cover, since they patch the same eviction sweep:
PR Liveness signal it adds Policy #48730 running terminals renew while any exist #50499 pending forms + permissions renew, whole location; explicit Stop still works #51827 live child processes skip expired entries with live children #51583 per-session progress (model output, tool updates) keeps a 1h bound for unanswered questions Two things worth deciding together:
- Policy for pending human waits. fix(core): preserve progressing sessions during location cleanup #51583's 1h bound for unanswered questions would still kill the headline case here — a question legitimately waiting on a human. The wait I measured before eviction was 5462.9s (~91 min; timeline in my earlier comment). fix(core): preserve pending human waits during idle cleanup #50499's shape (renew while a form/permission is pending, explicit Stop still interrupts) fits the reported failure better.
- These four patch the same ~5 lines with four different probes/registries. If you'd rather have one predicate, I'm happy to help consolidate (crediting everyone) — but that's a maintainer call, not something I'll do unilaterally.
Also worth knowing: branch
devno longer haslocation-activity.ts— it's back to the plain LayerMap idle TTL (location-services.ts, hardcoded"60 minutes") with no interrupt reason at all (SessionExecution.interrupton dev can't even express one). If that's unintentional, the next v2→dev merge drops this whole fix. Happy to file a separate issue if useful.Reacted by Ron SaksonovRelated reproduction from a different victim path — worth linking so the TTL is recognised as a broad defect, not just the shell-kill case.
On v2.0.22 a location was evicted at the 60-min idle TTL while a session in it was idle after dispatching a
background: trueshell job. The eviction didn't just interrupt a parked owner: it tore down the location-scoped plugin graph, which is where the background completion watcher is forked (ShellTool.notifyWhenDone,Effect.forkIn(scope, …),tool/plugin/shell.ts:154-186). The watcher fiber was interrupted beforesessions.synthetic(...)ran, so the completion notification was never emitted at all (verified: nosyntheticrow insession_messagefor that shellID;session_inbox/session_pendingempty) and the session was never resumed.Agree with the config ask here: the TTL is hardcoded (
location-activity.ts:25,options.timeToLive ?? "60 minutes") and no env/flag exists. A configurable or even merely higher default would also prevent this notification-loss case, so raising this alongside the parked-session case may help argue for the knob.Two findings from T3 Code on 2.0.24, in case they help:
- Eviction also drops MCP servers added at runtime (
PUT /api/experimental/mcp/:server): they live in the Location's in-memory overrides. T3 registers one per thread and never re-adds it, so after an idle hour the thread silently has no T3 tools until T3 restarts (no catalog notice either). Verified in a VM withPOST /api/location/reload, which tears down the same way. - A plugin stopgap for parked asks has to go through the HTTP API:
ctx.session.update(...)from plugin code publishes without the Location envelope (the promise adapter's runtime has noLocation.Service), so it renews nothing.PATCH /api/session/:idagainst the server's own port does. Verified with a stepped clock: without it the parked question/permission was cancelled at +75 min; with a periodic PATCH it stayed pending through +100 min.
- Eviction also drops MCP servers added at runtime (
Raised discussion internally about how best to approach the eviction logic, clearly there are some less than ideal states currently
Reacted by Khoa Huynh
Description
A session that is still running is interrupted when its Location expires after 60 minutes without a durable session event. Closing the web UI/browser does not stop a run by itself, but a run that is parked (e.g. on a
question) or otherwise silent produces no durable events, so the Location's deadline passes and the active execution is interrupted.How it works on 2.0.8 (
packages/core/src/location-activity.ts):LocationActivitykeeps a per-Location 60-minute deadline and renews it only on durable session events that carry alocation(Schema.is(SessionEvent.Durable)in itsbus.listen). Plain HTTP requests, includingGET /api/session/{id}, do NOT renew it.execution.interrupt(session.id, { reason: "inactivity", awaitSettlement: true })) before invalidating it, which is why the run ends asStep interrupted. Pending questions/forms are rejected by the Location finalizers (packages/core/src/question.ts,location-lifecycle.ts).Observed on a run that was parked:
The same session's last durable events land on the same second:
Two problems:
reason: "inactivity") even when the run is only parked on user input or waiting on a subagent.LocationActivity(itslayer({ timeToLive })defaults to"60 minutes"and the node is built withlayer()); there is no config, env var, or CLI flag to change it, and no way to opt out of interrupting running work.Related: #48691 (terminals/shells killed by the same 60-minute location eviction), #44471 (pending question form lost when a location is evicted/reconnect).
Plugins
usage, ci-wake, harness-review (local, server-side). No desktop terminal/PTY involved.
OpenCode version
2.0.8 (latest, standalone
opencode serve)Steps to reproduce
opencode serveon an always-on host and open the web UI in a browser on another machine.questiontool, or wait on a stuck subagent).location services evictedfor the directory and the run ends withStep interrupted(idle outcome=interrupted).In the observed incident the eviction timestamp and the
idle outcome=interruptedevent fall on the same second, 60 minutes after the Location was last renewed by a durable event.Screenshot and/or share link
No response
Operating System
macOS 26.6.2 (Darwin 25.6.0, arm64) on the server; client is a MacBook browser
Terminal
n/a — headless
opencode serveplus the web UI in a browser