Skip to content

local/unknown models wedge on context overflow: 200k default window + system-prompt-blind estimate + substring-only detection #203

Description

@justrach

Symptom

An unknown or locally-hosted model (LM Studio / mlx / any non-catalogued
base_url) can hit a hard, unrecoverable context wedge. The gaps chain:

  1. G8 - contextFor() falls back to default_context = 200_000 for any model
    not in the baked/overlay tables, so the window is a guess and every downstream
    cap/threshold is mis-sized. src/pricing.zig:303, src/provider.zig:150.
  2. G4 - the pre-send gate uses fullInputEstimateTokens(), which serializes
    only self.messages and omits the system prompt + tool schemas, so it
    under-counts and the gate under-fires. src/agent_request.zig:614-620.
  3. G2 - when the backend rejects, isContextOverflow() matches ~6 English
    substrings with no structured error.code / HTTP-status check, so a
    local provider whose message differs never triggers recovery and the meter is
    never pinned. src/agent_request.zig:81-92.

Result: bad window guess -> under-count -> gate never fires -> backend rejects ->
needle miss -> no recovery -> hard wedge (only manual /compact escapes).

Design note (from openai/codex research)

openai/codex has no magic for unknown local windows either - its fallback is a
single hardcoded 272k constant. Its real fix is escape hatches: a user config
override (model_context_window, clamped to max) plus a fetched/cached per-model
catalog (list_models). Note: OpenAI-compatible /v1/models for local servers
usually returns only id/object/owned_by, not a context window, so a probe
often falls back anyway - treat probing as a spike, not the plan.

Proposed fix (priority order)

  1. Window override (root fix, do first): a --context flag / config field that
    sets Provider.context directly, bypassing the 200k floor. Cascades to correct
    compactAt (G9), perOutputCap (G1), and the pre-send gate.
  2. G4 estimate baseline: add the serialized system-prompt + tool-schema byte
    count (or a fixed baseline constant, cf. codex BASELINE_TOKENS = 12000) to
    fullInputEstimateTokens before the /4, so the lower bound fires before the
    wall.
  3. G2 structured detection (backstop): in isContextOverflow, first inspect the
    parsed error object's code field and the HTTP status (400/413), then keep the
    substrings as a last-resort fallback. This is the one place we go beyond
    codex, which abandoned the non-streaming 400 path that is exactly the local
    wedge surface.
  4. (spike, optional) opportunistic metadata probe when a local base_url model
    isn't in the table; accept only if a window field is actually present.

Refs

#193, #192, #165

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions