Repository navigation
[Bug]: agent turns stop silently mid-task: the planned tool call is swallowed, provider returns a clean finish_reason=stop and the billed usage exceeds the delivered content #17500
Description
Activity
- addedbugSomething is broken or behaving incorrectly.Something is broken or behaving incorrectly.needs-triageIssue needs maintainer review and initial categorization.Issue needs maintainer review and initial categorization.
on Oct 9, 2026 Note
Grok responding on behalf of Julius.
Thanks for the 2x2 matrix and the token accounting. That makes the T3 build the variable, and it points at the request T3 hands Pi rather than at Pi itself. I haven't reproduced it; the notes below come from reading the code.
What changed between the two nightlies
There are 166 commits between
v0.0.46-nightly.20261005.2667andv0.0.46-nightly.20261009.2873. Most of the Pi changes first ship together inv0.0.46-nightly.20261008.2849:- fix(server): Pi discovers optional T3 tools on demand #17220 (Pi discovers optional T3 tools on demand). Top suspect. On Pi versions that expose
registerMcpServer(both 1.0.4 and 1.1.0 should qualify, which would fit your matrix), every T3 MCP tool exceptorchestrator_capabilities,delegate_taskandtask_statusis now registered asexposure: "deferred"(L309-L331). It also adds a hiddenmcp__t3_code__alias for each tool (L320-L325) and adds Pi's builtintool_searchto the active set (L370-L386). The same PR dropped the per-toolpromptSnippet/promptGuidelineslines, so the system prompt changed too. It closed [Bug]: Pi sessions declare every t3-code tool on every turn #16651, which asked to stop declaring every tool on every turn. - fix(server): Pi extension wakes get an owned continuation turn #17214 (an owned continuation turn for Pi extension wakes) changes how T3 settles turns, not what gets sent to the model, so it seems less likely to explain a model-side
finish_reason: "stop". - refactor(provider-pi): move Pi into its own provider package #17302 (Pi moved into
packages/provider-pi) looks like a move, but it landed in the same nightly. - In later nightlies, fix(server): Pi loads every selected skill without losing prompt text #17194 (Pi skill loading,
20261009.2861) and fix(server): reconcile Pi native session rewinds #13839 (Pi session rewinds,20261009.2873) touch Pi too. They seem less related to the tool declarations.
A possible mechanism (unverified)
After #17220, a request may declare only a few T3 tools plus
tool_search, while the conversation history (and the model's plan) still refers to tools likemcp__t3-code__t3_thread_sendthat aren't declared in that request. If the model tries to call one of those tools directly instead of searching first, an OpenAI-compatible gateway or backend that validates or constrains tool calls againsttoolsmight drop the call and return a cleanstop. Billed output tokens would then exceed the delivered content by about the size of the call, as you saw. Your mock runs sent all 34 schemas declared up front, which could explain why they don't reproduce it.Detection
As far as I can tell, neither T3 nor Pi can tell this apart from a normal finished turn, since the stream ends with
stopand no tool call. Comparing usage against delivered tokens might be a possible heuristic, but nothing does that today.Could you help narrow it down?
- Try
v0.0.46-nightly.20261008.2833(the last build before the batch) andv0.0.46-nightly.20261008.2849(the first build with fix(server): Pi discovers optional T3 tools on demand #17220, fix(server): Pi extension wakes get an owned continuation turn #17214 and refactor(provider-pi): move Pi into its own provider package #17302). If 2833 is clean and 2849 fails, that narrows it to those three PRs. - On a failing build, without sharing any content, could you check whether the swallowed call targets a T3 tool other than the three direct ones, and whether the request's
toolsarray is much shorter than 34? - If Pi lets you disable or replace
tool_search, the extension falls back to declaring every tool directly (L370-L380). That might work as a temporary workaround, but I haven't tested it.
- fix(server): Pi discovers optional T3 tools on demand #17220 (Pi discovers optional T3 tools on demand). Top suspect. On Pi versions that expose
- addedvia-triageFiled through npx t3 triageFiled through npx t3 triageand removedneeds-triageIssue needs maintainer review and initial categorization.Issue needs maintainer review and initial categorization.
on Oct 9, 2026 i did it and yeah
v0.0.46-nightly.20261008.2833is clean,v0.0.46-nightly.20261008.2849reproduces it. so the regression is in whatever landed in 2849 and #17220 (deferred tool exposure) matches the logic exactlyto answer to your questions
which tools get swallowed
every call we could identify was targeting a deferred t3 tool (
t3_thread_send,t3_thread_read). never one of the three direct ones (orchestrator_capabilities,delegate_task,task_status)what the declared tools array actually looks like on the failing build
i ran a tiny local logging proxy in front of our gateway and captured the exact array a fresh session sends
deferred mode (default) forced direct declared tools 22 99 t3 mcp tools declared 3 80 tool_searchdeclared declared (shadow) so with the default setup the model can discover and call
t3_thread_sendetc. throughtool_searchwhile those names are absent from the declared array. our gateway appears to drop the tool call server side when the name is not declared, then ends the stream with a cleanfinish_reason=stopwhile usage is billed as if it completed. that matches every specimen we have (billed output tokens exceed delivered content, and the missing chunk is the size of the planned tool call)workaround
a pi extension registering a tool named
tool_searchreplaces the builtin (same idea the docs mention for codemode), which makes the t3 bridge treat discovery as unavailable and fall back to declaring everything directly. tested on 2849, same build, same gateway- before the shadow, a call to
t3_thread_sendwas being swallowed repeatedly, including in my own agent session mid-conversation - after installing the shadow, the same call went through (
finish_reason=tool_calls, then a clean final stop) and no swallow happened again
so anyone on the nightly can pin that extension as a temporary fix until this is reverted or gated
- before the shadow, a call to
Before submitting
Area
apps/server
Steps to reproduce
Expected behavior
a turn either delivers its tool call (and it runs) or surfaces a provider error. never a clean, silent end of turn with the announced action missing
Actual behavior
the provider response ends with finish_reason="stop" and no tool_call delta, while usage.output_tokens is larger than everything actually delivered (reasoning + text). the missing remainder matches the size of the planned tool call. the client honors the clean stop and marks the turn complete, so nothing retries and nothing reports an error.
measured specimen (raw pi session jsonl, 2026-10-09 12:38 UTC):
12 confirmed specimens across 2 days, same signature
bisect, full 2x2 matrix (same t3_thread_send prompt, same session state):
so the pi version has no effect, the t3 code build is the only variable that decides whether tool calls get swallowed from my testing
controls: 163 tool-heavy mock runs through the same gateway (same 34 tool schemas, same prompt shapes, synthetic MCP server) never reproduce, so the trigger needs whatever request shape the newer build emits
i rolled back to 0.0.46-nightly.20261005.2667 and it works again with both Pi 1.0.4 and 1.1.0
Impact
Major degradation or frequent failure
Version or commit
0.0.46-nightly.20261009.2873
Environment
macOS 27.0 (apple silicon), t3 code nightly desktop app, provider: pi (pi 1.1.0 and 1.0.4 both tested), custom openai-compatible provider behind a private gateway
Logs or stack traces
not providing any logs bc my provider is a private gateway and everything going through it is confidential, so i can't share request or response content. it reproduces very easily though with any long tool-heavy threads on 20261009.2873 hit it within minutes, and the exact same prompt on 20261005.2667 is clean, so you don't need my logs to repro it on your sideScreenshots, recordings, or supporting files
No response
Workaround
rolling back to a previous version like 0.0.46-nightly.20261005.2667