Skip to content

[Bug]: agent turns stop silently mid-task: the planned tool call is swallowed, provider returns a clean finish_reason=stop and the billed usage exceeds the delivered content #17500

Description

@eliottscherrer

Before submitting

  • I searched existing issues and did not find a duplicate.
  • I included enough detail to reproduce or investigate the problem.

Area

apps/server

Steps to reproduce

  1. register a custom OpenAI-compatible provider (private gateway, standard /chat/completions API), as a pi provider instance
  2. on 0.0.46-nightly.20261009.2873, run an agent thread on a real task (multi-step work, full T3 MCP tool list)
  3. turns die mid-task: the assistant announces its next action, then the turn ends cleanly, no tool call runs, no error shown, the thread shows completed with half the work done
  4. same session, same prompt, same state on 0.0.46-nightly.20261005.2667: the exact same call that was swallowed 5/5 times succeeds on the first attempt after rollback

Expected behavior

a turn either delivers its tool call (and it runs) or surfaces a provider error. never a clean, silent end of turn with the announced action missing

Actual behavior

the provider response ends with finish_reason="stop" and no tool_call delta, while usage.output_tokens is larger than everything actually delivered (reasoning + text). the missing remainder matches the size of the planned tool call. the client honors the clean stop and marks the turn complete, so nothing retries and nothing reports an error.

measured specimen (raw pi session jsonl, 2026-10-09 12:38 UTC):

  • billed: output 2852 tokens (input 224512)
  • delivered: 2757 reasoning tokens + 128 chars of visible text (~35 tokens)
  • remainder ~60 tokens: the planned tool call, never emitted
  • stopReason "stop", zero toolCalls in the message, no error

12 confirmed specimens across 2 days, same signature

bisect, full 2x2 matrix (same t3_thread_send prompt, same session state):

t3 build pi 1.0.4 pi 1.1.0
20261009.2873 dead dead (5/5 attempts swallowed)
20261005.2667 clean clean (verified on a fresh thread after flipping pi back to 1.1.0, the exact same send succeeds first try)

so the pi version has no effect, the t3 code build is the only variable that decides whether tool calls get swallowed from my testing

controls: 163 tool-heavy mock runs through the same gateway (same 34 tool schemas, same prompt shapes, synthetic MCP server) never reproduce, so the trigger needs whatever request shape the newer build emits

i rolled back to 0.0.46-nightly.20261005.2667 and it works again with both Pi 1.0.4 and 1.1.0

Impact

Major degradation or frequent failure

Version or commit

0.0.46-nightly.20261009.2873

Environment

macOS 27.0 (apple silicon), t3 code nightly desktop app, provider: pi (pi 1.1.0 and 1.0.4 both tested), custom openai-compatible provider behind a private gateway

Logs or stack traces

not providing any logs bc my provider is a private gateway and everything going through it is confidential, so i can't share request or response content. it reproduces very easily though with any long tool-heavy threads on 20261009.2873 hit it within minutes, and the exact same prompt on 20261005.2667 is clean, so you don't need my logs to repro it on your side

Screenshots, recordings, or supporting files

No response

Workaround

rolling back to a previous version like 0.0.46-nightly.20261005.2667

Activity

  1. added
    bugSomething is broken or behaving incorrectly.
    needs-triageIssue needs maintainer review and initial categorization.
    on Oct 9, 2026
  2. juliusmarminge commented on Oct 9, 2026

    @juliusmarminge
    Member

    Note

    Grok responding on behalf of Julius.

    Thanks for the 2x2 matrix and the token accounting. That makes the T3 build the variable, and it points at the request T3 hands Pi rather than at Pi itself. I haven't reproduced it; the notes below come from reading the code.

    What changed between the two nightlies

    There are 166 commits between v0.0.46-nightly.20261005.2667 and v0.0.46-nightly.20261009.2873. Most of the Pi changes first ship together in v0.0.46-nightly.20261008.2849:

    A possible mechanism (unverified)

    After #17220, a request may declare only a few T3 tools plus tool_search, while the conversation history (and the model's plan) still refers to tools like mcp__t3-code__t3_thread_send that aren't declared in that request. If the model tries to call one of those tools directly instead of searching first, an OpenAI-compatible gateway or backend that validates or constrains tool calls against tools might drop the call and return a clean stop. Billed output tokens would then exceed the delivered content by about the size of the call, as you saw. Your mock runs sent all 34 schemas declared up front, which could explain why they don't reproduce it.

    Detection

    As far as I can tell, neither T3 nor Pi can tell this apart from a normal finished turn, since the stream ends with stop and no tool call. Comparing usage against delivered tokens might be a possible heuristic, but nothing does that today.

    Could you help narrow it down?

  3. added
    via-triageFiled through npx t3 triage
    and removed
    needs-triageIssue needs maintainer review and initial categorization.
    on Oct 9, 2026
  4. eliottscherrer commented on Oct 9, 2026

    @eliottscherrer
    Author

    i did it and yeah v0.0.46-nightly.20261008.2833 is clean, v0.0.46-nightly.20261008.2849 reproduces it. so the regression is in whatever landed in 2849 and #17220 (deferred tool exposure) matches the logic exactly

    to answer to your questions

    which tools get swallowed

    every call we could identify was targeting a deferred t3 tool (t3_thread_send, t3_thread_read). never one of the three direct ones (orchestrator_capabilities, delegate_task, task_status)

    what the declared tools array actually looks like on the failing build

    i ran a tiny local logging proxy in front of our gateway and captured the exact array a fresh session sends

    deferred mode (default) forced direct
    declared tools 22 99
    t3 mcp tools declared 3 80
    tool_search declared declared (shadow)

    so with the default setup the model can discover and call t3_thread_send etc. through tool_search while those names are absent from the declared array. our gateway appears to drop the tool call server side when the name is not declared, then ends the stream with a clean finish_reason=stop while usage is billed as if it completed. that matches every specimen we have (billed output tokens exceed delivered content, and the missing chunk is the size of the planned tool call)

    workaround

    a pi extension registering a tool named tool_search replaces the builtin (same idea the docs mention for codemode), which makes the t3 bridge treat discovery as unavailable and fall back to declaring everything directly. tested on 2849, same build, same gateway

    • before the shadow, a call to t3_thread_send was being swallowed repeatedly, including in my own agent session mid-conversation
    • after installing the shadow, the same call went through (finish_reason=tool_calls, then a clean final stop) and no swallow happened again

    so anyone on the nightly can pin that extension as a temporary fix until this is reverted or gated

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething is broken or behaving incorrectly.via-triageFiled through npx t3 triage

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions