Skip to content

Exercise the OpenAI-compatible (BYO-key) agent end-to-end against a free hosted provider, and fix what breaks #2189

Description

@malibio

Summary

Exercise the OpenAI-compatible (BYO-key) agent path end-to-end against a real hosted provider, driven from the UI the way a user would drive it: add a key in Settings, pick that endpoint in an AI Chat, and run agentic requests ("create a task to XYZ", "create an author schema"). Fix what breaks.

The transport, discovery, settings storage and model selector all ship today. What has never been exercised is a real remote provider over the public internet with a real API key — every measurement to date was against a local Ollama at http://127.0.0.1:11434/v1. That difference is where the untested surface lives: HTTPS, bearer auth, provider-specific /models and /chat/completions dialects, rate limits, and latency an order of magnitude above localhost.

Use a free tier so this stays reproducible without spend — Google AI Studio (https://aistudio.google.com) issues a free key and exposes an OpenAI-compatible endpoint:

base_url: https://generativelanguage.googleapis.com/v1beta/openai
model:    gemini-flash-latest   (confirm the current free-tier id in AI Studio)
api_key:  <AI Studio key>

Any free OpenAI-compatible provider is acceptable if that one proves unworkable; record which was used and why.

Why now

grep -ri "gemini\|googleapis\|aistudio" over the Rust/Svelte source returns nothing in the inference path — the only hits are the packages/skill CLI shims (Gemini CLI as a host, unrelated) and a core_schemas.rs enum value. No hosted provider has ever touched OpenAiCompatInferenceEngine.

Scope

1. Configure from Settings, verify discovery

  • Add the AI Studio config via Settings → Model Manager (name, base URL, API key, model). Do not hand-edit daemon.toml — the point is to test the path a user takes.
  • Confirm the endpoint's models appear in the remote-models list. discover_models does GET {base_url}/models with a 5s timeout and expects {"data":[{"id":...}]}; verify the provider answers in that shape and within that budget over the public internet.
  • Confirm the API key is stored and sent as Authorization: Bearer, and that it is not written to logs or the prompt dump.

2. Select the endpoint in an AI Chat

  • In an AI Chat node, open the model selector and pick the AI Studio entry.
  • Confirm the selection persists on the node (provider: "openai-compat", openai-compat:<uuid>:<model>) and survives a reload.

3. Run agentic requests

At minimum, and via the chat UI rather than a test harness:

  • create a task to XYZ → expect a create_node call producing a real task node
  • create an author schema → expect create_schema with sensible fields
  • one query turn (e.g. find my notes about X) → expect search_nodes
  • one update turn → expect the resolve_query → update_node handoff

Record for each: tools called, whether the node/schema actually landed in the DB, and the assistant's final message.

4. Fix what breaks

Transport, auth, discovery, parsing, streaming, tool-call extraction, error surfacing — fix in this issue. Two known-shaky areas to look at first:

  • The routing probe (packages/agent/src/local_agent/routing_probe.rs). It runs one synthetic routed turn per model load and disables Stage-2 candidate injection for models that fail. It was calibrated entirely against local Ollama; against a hosted provider it now costs a billed network round-trip on every model load, and a rate-limit or transport error at probe time may be indistinguishable from a genuine "model can't route" verdict — which would silently and permanently disable routing for a model that routes fine. Verify which verdict AI Studio gets and that it is earned, not an artifact.
  • EXCLUDED_MODEL_FRAGMENTS in tests/live_openai_compat_routing.rs is ["qwen", "ornith"]. Per ADR-056 those models are removed and are not evidence; confirm a Gemini id isn't caught by the filter and decide whether the live matrix should now cover a hosted arm.

Success criteria — read this carefully

A working transport with poor agent behavior is a PASS, not a failure.

If the request completes, tools are invoked, and results land in the database — but the agent chooses the wrong tool, over-asks, returns prose instead of acting, or produces a badly-shaped schema — then this issue succeeded and the prompting deficiency is a separate new issue, filed with the concrete transcripts as evidence. Do not fix prompting here.

That split matters because the two failure modes have different owners and different evidence. Prior work established repeatedly that prompt-quality problems are only diagnosable from a dumped prompt (NODESPACE_PROMPT_DUMP — the openai-compat path has its own dumper at openai_compat_prompt_dump.rs, tagged "engine": "openai_compat"), and that tuning prose without one is guesswork.

  • AI Studio config added through Settings; models discovered
  • Endpoint selectable in AI Chat; selection persists
  • All four agentic requests run to completion with tool calls
  • Transport/auth/parsing defects found are fixed here
  • Routing-probe verdict for the hosted model verified as earned
  • API key confirmed absent from logs and prompt dumps
  • If agent behavior is weak-but-working: new prompting issue filed with transcripts, linked here
  • Prompt dump captured for at least one failing/weak turn, attached to whichever issue owns it

Out of scope

  • Prompt/guidance tuning (→ new issue per above)
  • Shipping AI Studio as a curated preset — this is a BYO-key config like any other
  • The native GGUF path; ADR-056 keeps it locked to one model
  • Non-OpenAI-compatible providers (Anthropic/Gemini native wire formats)

Relevant code

  • packages/agent/src/local_agent/openai_compat_inference.rs — the engine
  • packages/agent/src/local_agent/openai_compat_discovery.rs — GET /models, 5s timeout
  • packages/agent/src/local_agent/routing_probe.rs — per-model routing verdict
  • packages/daemon/src/services/settings_service.rs — config storage, probe-verdict carry-forward
  • packages/desktop-app/src/lib/components/settings/model-manager.svelte — the settings form
  • packages/desktop-app/src/lib/components/viewers/ai-chat-model-selector.svelte — the chat picker
  • packages/agent/tests/live_openai_compat_smoke.rs, live_openai_compat_routing.rs — existing local-Ollama coverage

Activity

  1. mstomar125 commented on Aug 21, 2026

    @mstomar125
    Contributor

    Picked this up but it's blocked at step 1: this issue needs a real API key from a free hosted OpenAI-compatible provider (e.g. Google AI Studio), which requires an interactive browser/account signup this environment doesn't have (standing constraint — no browser/Google account access on this remote box, deferred to local-machine sessions in prior work). Checked for an existing stored credential in ~/nodespace/.env and local NodeSpace config — none found. Leaving this on the board as Todo rather than assigning it further autonomous work; needs either a human to provision the key and hand it off, or to be picked up from a machine with browser access.

  2. malibio commented on Aug 21, 2026

    @malibio
    CollaboratorAuthor

    @mstomar125 — clarification on the blocker: create the key yourself. AI Studio keys are free and self-serve, so there is no need to wait for one to be handed over.

    No key should be passed around, and none should go into a .toml or .env — that is not the flow under test. This issue is specifically about the GUI path a user would take: open Settings, add the provider with your own key, pick that endpoint as the AI Chat model, then run the agentic requests. Entering it through Settings is part of what is being exercised, so checking ~/nodespace/.env or the daemon config was looking in the wrong place by design.

    If browser access on that box is a hard constraint, this one needs to move to a machine that has it rather than be worked around — provisioning is a normal one-time setup step here, not a credential handoff.

    Two things worth knowing before you start, both of which will otherwise look like new bugs:

    #2198 is on this exact path and is yours. "OpenAI-compat engine drops assistant tool_calls from replayed history" — any multi-turn agentic request replays history, so the model loses the record of what it just did. Worth fixing first; otherwise you will rediscover it by hand here.

    The routing probe will fire on a new served model. Stage-2 candidate injection suppresses tool-calling outright on some served models — measured on mistral:7b, where any injected block suppressed, regardless of content (#1828/#1829). A one-time probe per (base_url, model) caches the verdict and sets routing_disabled, so a Gemini endpoint may silently run with Stage-2 injection off — a different pipeline than the native path. Read the probe verdict out of the daemon log before drawing conclusions about behavior, or a routing difference will read as a model difference.

  3. malibio commented on Aug 21, 2026

    @malibio
    CollaboratorAuthor

    @mstomar125 — assigning this to you. To be explicit about scope, since "exercise it end-to-end" can be read narrowly:

    This is a test-and-fix task, not just a test task. Drive the BYO-key path through the GUI as a user would, then act on whatever it turns up:

    • Fix what is in scope — anything on the OpenAI-compat path itself (transport, auth, /models and /chat/completions dialect handling, settings persistence, model selection, the agentic turn). Fix it here.
    • File what is not — a finding that is really a separate defect, or that needs its own investigation or design decision, gets its own issue with a repro. Do not bundle unrelated fixes into this PR.
    • Record what you could not reach. If a provider quirk blocks a path, say which and why rather than leaving it silent. A negative result is a result.

    Create the AI Studio key yourself — free and self-serve. Enter it through Settings in the running app, not a .toml or .env: the settings UI is part of what is under test, so bypassing it skips the surface this issue exists to exercise. No key gets handed over or committed.

    Two known hazards on this exact path, so they read as expected rather than as new discoveries:

    #2198 is yours and sits directly under this. The OpenAI-compat engine drops assistant tool_calls from replayed history, so any multi-turn agentic request loses the record of what it just did. Worth fixing first — otherwise you will rediscover it by hand here, and it will contaminate every multi-turn result you record.

    The routing probe fires on a new served model. Per #1828/#1829, Stage-2 candidate injection suppresses tool-calling outright on some served models — measured on mistral:7b, where any injected block suppressed regardless of content. A one-time probe per (base_url, model) caches the verdict and sets routing_disabled, so a Gemini endpoint may silently run with Stage-2 injection off — a materially different pipeline from the native path. Read the probe verdict out of the daemon log before drawing conclusions, or a routing difference will look like a model difference.

    One methodological note, learned the hard way in #2188: if something looks wrong, dump the prompt before theorising about it. NODESPACE_PROMPT_DUMP=<path> captures the exact request at the chokepoint, and the OpenAI-compat path has its own capture point (openai_compat_prompt_dump.rs) writing to the same file with "engine": "openai_compat". Three hypotheses in that PR were refuted by a dump that took seconds to run.

  4. malibio commented on Aug 21, 2026

    @malibio
    CollaboratorAuthor

    Starting on this. One blocker: exercising this end-to-end requires a real API key for a free hosted OpenAI-compatible provider (Google AI Studio per the issue, or another free provider). I don't have one and can't generate one myself — I'll need it supplied (e.g. pasted into Settings → Model Manager directly by whoever has an AI Studio account, or handed to me to enter).

    Proceeding in the meantime with the startup sequence and static investigation of the two known-shaky areas (routing_probe.rs verdict logic, EXCLUDED_MODEL_FRAGMENTS in live_openai_compat_routing.rs) so I'm ready to move fast once a key is available.

  5. malibio commented on Aug 21, 2026

    @malibio
    CollaboratorAuthor

    Baseline: Frontend 4794 passed (227 test files), Duration 79.23s. Rust baseline requires a warm cargo cache — will run as part of test:all before PR.

  6. self-assigned this
    on Aug 21, 2026
  7. malibio commented on Aug 21, 2026

    @malibio
    CollaboratorAuthor

    Correction to my earlier comment: I do have a self-serve path to an AI Studio key — I misread the constraint as 'need a key handed to me' when it's actually 'create one myself, enter it through Settings.' The actual blocker is that key creation needs interactive browser/Google-account access, which I don't have as a CLI agent. I'll flag this to Michael directly rather than block silently.

    In the meantime: I'm picking up #2198 first (assistant tool_calls dropped from replayed history), since it directly contaminates any multi-turn agentic test here and per Michael's guidance above should land before the walkthrough. Will return to this issue once that's merged and a key is available.

  8. malibio commented on Aug 22, 2026

    @malibio
    CollaboratorAuthor

    @mstomar125 — OpenRouter is a second option for the free key, if AI Studio proves awkward. Either is fine; the issue only needs a real hosted OpenAI-compatible endpoint reached over the public internet.

    base_url: https://openrouter.ai/api/v1
    api_key:  <your OpenRouter key>
    model:    <one of the :free ids below>
    

    Verified against their live API today (2026-08-22): 421 models, 18 with a :free suffix, and 17 of those 18 advertise tools in supported_parameters — so tool calling, which is the entire point of this test, is available on the free tier rather than paywalled.

    A few free ids that list tool support, with their context windows:

    model ctx
    z-ai/glm-5.2:free 256k
    cohere/north-mini-code:free 256k
    thinkingmachines/inkling:free 262k
    poolside/laguna-s-2.1:free 262k
    nvidia/nemotron-3.5-lightning:free 1M
    liquid/lfm-2.5-2.6b:free 64k

    Re-check the list at GET https://openrouter.ai/api/v1/models before choosing — free-tier availability rotates, and supported_parameters is the field that tells you whether a given model will accept a tool call at all.

    Why this may be the better choice than AI Studio here: OpenRouter is a strict OpenAI-compatible implementation, which is precisely the surface this issue exists to exercise. Local Ollama accepts malformed payloads silently — that laxness is why the tool_calls serialization bug (#2198, fixed in PR #2236) survived unnoticed while every prior measurement ran against it. A strict endpoint fails loudly on a contract violation, which is the signal we want.

    Being able to swap model across several providers behind one key and one base_url is also useful for the routing-probe concern noted above: if one served model comes back routing_disabled you can try another without reconfiguring anything, and record the verdict per model.

    Whichever you pick, record which provider and model you used and why — the result is not portable across endpoints, so the next person needs to know what was actually exercised.

    Still applies: create the key yourself, enter it through Settings in the running app, no key in a .toml, .env, or committed anywhere.

  9. mstomar125 commented on Sep 14, 2026

    @mstomar125
    Contributor

    Moved to Backlog — deprioritized for now, not being worked right now. Still needs a real OpenRouter or Google AI Studio key (account creation isn't something an agent can self-serve) before the end-to-end agentic test can run. Instructions and free model recommendations are already in this thread, ready to go whenever this comes back up.

  10. removed their assignment
    on Sep 28, 2026
  11. malibio commented on Oct 5, 2026

    @malibio
    CollaboratorAuthor

    One more case for this pass, from the pinned-chat work (PR #3575): a play's Edit chat opens on a message the system writes, so its first turn sends the model system, assistant, system, user: a conversation whose first non-system message is the assistant's. That shape is verified on the native Gemma 4 E4B template (nlp-engine tests/it/assistant_first_conversation.rs) and has not been tried against an OpenAI-compatible endpoint, where a server strict about role order could reject it. To check: open Edit on a play with an OpenAI-compatible model selected and send the first reply.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions