Skip to content

mcp.enabled is store-authoritative but mcp.network is file-authoritative, and the disagreement logs as success #2389

Description

@erikdarlingdata

mcp.enabled has two authorities that can disagree, the store silently wins, and the only evidence is one INFO line after a success message. Setting the config file alone looks like it worked for about five seconds.

Hit while wiring up the MCP on prod-sql-use1-pgmonitor-01.

What it looks like

darling.json had "mcp": { "enabled": true, "port": 5152, "network": { ... } }. The service read it, started the server, bound the LAN address, and logged success:

22:58:51 [INFO] [DarlingMcpHostService] Starting MCP server on http://10.0.0.25:5152
                (LAN-exposed to 10.0.0.0/16 behind a bearer token + in-app CIDR; loopback also bound)
22:58:56 [INFO] [DarlingMcpHostService] MCP server disabled via the control plane - stopping (no restart needed)

Five seconds apart. config.config_service.mcp_enabled was false, and the supervisor loop reconciles to it:

var enabled = published?.Enabled ?? config.Mcp.Enabled;

The store wins whenever a published row exists, which on any seeded store it always does. The file value is effectively dead.

Why this is worth a diagnostic rather than a doc note

The failure presents as success. An operator edits the config, restarts, greps the log, finds Starting MCP server on http://… and stops reading — that line is true, and the server really did bind. The contradiction is five seconds later at the same INFO level with no error, no warning, and nothing tying it back to the file they just edited.

It also inverts the usual mental model on this box family. darling.json is where MCP network exposure is configured — listen, allowFrom, encryptedToken are file-only and have no store equivalent. So mcp.network is file-authoritative while mcp.enabled sitting right beside it is store-authoritative. Same object, two different owners, no indication which is which.

This is the same class as the darling.json seeding trap already documented elsewhere ("the store, not the config file, is authoritative for servers") — but that one at least fails by doing nothing visible, whereas this one actively reports the opposite of what happens.

Suggested fix

Warn on the disagreement, not just the outcome. When a published Enabled overrides a differing config.Mcp.Enabled, say so once at the point of override, naming both sides:

MCP is enabled in darling.json but DISABLED in the control plane (config.config_service.mcp_enabled);
the store wins. Set it in the store, or the file value will keep being ignored.

Cheap, fires only on the mismatch, and lands exactly where someone is looking. Same treatment would suit port (published?.Port ?? config.Mcp.Port), which has the identical shape and would produce a server on an unexpected port with no explanation.

Worth considering whether the success log should be deferred until after the supervisor's first reconcile, so "Starting MCP server on …" is not printed for a server that is about to be stopped.

Activity

  1. added 2 commits that reference this issue on Aug 21, 2026
  2. erikdarlingdata commented on Aug 21, 2026

    @erikdarlingdata
    OwnerAuthor

    Fixed on dev by #2411.

    The resolution table, established from source before anything was changed:

    field darling.json store column wins at runtime
    enabled mcp.enabled config_service.mcp_enabled store
    port mcp.port config_service.mcp_port store
    network.listen mcp.network.listen — file, restart-only
    network.allowFrom mcp.network.allowFrom — file, restart-only
    network.encryptedToken mcp.network.* — file, restart-only

    web is the byte-identical twin. No store column exists for any network field, confirmed against the migrations.

    Two details change how the original report reads. The file is not a fallback that lost: SeedServiceRowAsync writes it once with ON CONFLICT DO NOTHING, so on a seeded store it is dead input — and the worker publishes the file values itself when the store read fails, so published is non-null in both the healthy and the store-down case. And the MCP host calls DarlingConfig.Load() itself and holds that instance for the process lifetime, so ApplyToConfig never reaches it. The two values genuinely coexist in one process.

    Making network store-authoritative was rejected, and DPAPI is only half the reason. The token is LocalMachine-scoped so it would be undecryptable on any other box — but the larger objection is that a bind address, CIDR and token in config_service would let a remote admin store connection re-point the listener onto a LAN interface behind its own credential. The split stays; it is now visible instead of silent.

    Resolution runs through one shared helper carrying provenance. The start line names which plane supplied each half of the bind and admits when it is provisional; a genuine disagreement warns at the point of override, naming both sides, the winner, the verb that changes it, and the network block's opposite ownership — only on mismatch, once per distinct state. Three CLI notes had the same defect in miniature, printing only when the file said disabled, which is backwards; they are now unconditional and about precedence.

    Worth recording: darling.sample.json already documented the split correctly on both blocks. The docs were right and the runtime was silent, so no documentation change would have prevented this — which is the general shape of it. A follow-up in the same family is filed as #2414: the firewall verbs name their rule from the file's port while the endpoint binds the store's, so changing the port in the Viewer yields an exposed endpoint with no firewall path.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions