Symptom
An unknown or locally-hosted model (LM Studio / mlx / any non-catalogued
base_url) can hit a hard, unrecoverable context wedge. The gaps chain:
- G8 -
contextFor() falls back to default_context = 200_000 for any model
not in the baked/overlay tables, so the window is a guess and every downstream
cap/threshold is mis-sized. src/pricing.zig:303, src/provider.zig:150.
- G4 - the pre-send gate uses
fullInputEstimateTokens(), which serializes
only self.messages and omits the system prompt + tool schemas, so it
under-counts and the gate under-fires. src/agent_request.zig:614-620.
- G2 - when the backend rejects,
isContextOverflow() matches ~6 English
substrings with no structured error.code / HTTP-status check, so a
local provider whose message differs never triggers recovery and the meter is
never pinned. src/agent_request.zig:81-92.
Result: bad window guess -> under-count -> gate never fires -> backend rejects ->
needle miss -> no recovery -> hard wedge (only manual /compact escapes).
Design note (from openai/codex research)
openai/codex has no magic for unknown local windows either - its fallback is a
single hardcoded 272k constant. Its real fix is escape hatches: a user config
override (model_context_window, clamped to max) plus a fetched/cached per-model
catalog (list_models). Note: OpenAI-compatible /v1/models for local servers
usually returns only id/object/owned_by, not a context window, so a probe
often falls back anyway - treat probing as a spike, not the plan.
Proposed fix (priority order)
- Window override (root fix, do first): a
--context flag / config field that
sets Provider.context directly, bypassing the 200k floor. Cascades to correct
compactAt (G9), perOutputCap (G1), and the pre-send gate.
- G4 estimate baseline: add the serialized system-prompt + tool-schema byte
count (or a fixed baseline constant, cf. codex BASELINE_TOKENS = 12000) to
fullInputEstimateTokens before the /4, so the lower bound fires before the
wall.
- G2 structured detection (backstop): in
isContextOverflow, first inspect the
parsed error object's code field and the HTTP status (400/413), then keep the
substrings as a last-resort fallback. This is the one place we go beyond
codex, which abandoned the non-streaming 400 path that is exactly the local
wedge surface.
- (spike, optional) opportunistic metadata probe when a local
base_url model
isn't in the table; accept only if a window field is actually present.
Refs
#193, #192, #165
Symptom
An unknown or locally-hosted model (LM Studio / mlx / any non-catalogued
base_url) can hit a hard, unrecoverable context wedge. The gaps chain:contextFor()falls back todefault_context = 200_000for any modelnot in the baked/overlay tables, so the window is a guess and every downstream
cap/threshold is mis-sized.
src/pricing.zig:303,src/provider.zig:150.fullInputEstimateTokens(), which serializesonly
self.messagesand omits the system prompt + tool schemas, so itunder-counts and the gate under-fires.
src/agent_request.zig:614-620.isContextOverflow()matches ~6 Englishsubstrings with no structured
error.code/ HTTP-status check, so alocal provider whose message differs never triggers recovery and the meter is
never pinned.
src/agent_request.zig:81-92.Result: bad window guess -> under-count -> gate never fires -> backend rejects ->
needle miss -> no recovery -> hard wedge (only manual
/compactescapes).Design note (from openai/codex research)
openai/codex has no magic for unknown local windows either - its fallback is a
single hardcoded 272k constant. Its real fix is escape hatches: a user config
override (
model_context_window, clamped to max) plus a fetched/cached per-modelcatalog (
list_models). Note: OpenAI-compatible/v1/modelsfor local serversusually returns only
id/object/owned_by, not a context window, so a probeoften falls back anyway - treat probing as a spike, not the plan.
Proposed fix (priority order)
--contextflag / config field thatsets
Provider.contextdirectly, bypassing the 200k floor. Cascades to correctcompactAt(G9),perOutputCap(G1), and the pre-send gate.count (or a fixed baseline constant, cf. codex
BASELINE_TOKENS = 12000) tofullInputEstimateTokensbefore the/4, so the lower bound fires before thewall.
isContextOverflow, first inspect theparsed error object's
codefield and the HTTP status (400/413), then keep thesubstrings as a last-resort fallback. This is the one place we go beyond
codex, which abandoned the non-streaming 400 path that is exactly the local
wedge surface.
base_urlmodelisn't in the table; accept only if a window field is actually present.
Refs
#193, #192, #165