Skip to content

Firewall residues: an upgrade opens the pre-change port for one cycle, and a disabled surface still gets a rule #2436

Description

@erikdarlingdata

Two residues from #2414/#2432, which fixed the firewall verbs to resolve their port through the same resolver the endpoint binds on. Both were disclosed rather than folded in, and both are now visible instead of silent — which is why neither is urgent.

1. An in-place upgrade of a moved-port box opens the wrong port for one cycle

install-darling.ps1 runs --configure-firewall at line 521 and Start-Service at line 534. The firewall is therefore configured before the service has started, which is correct on a fresh install — config_service is seeded from darling.json, so the file is the control plane's future answer and cannot be wrong.

It is not correct on an upgrade of a box whose port was later changed in the Viewer. There the store already holds a different port, the service is stopped, and --configure-firewall falls back to the file. So the rule lands on the old port for one cycle.

This was true before #2432 as well; what changed is that the fallback now says so, and the service's own start-up check WARNs the exact command to fix it. So it self-heals with one operator action rather than silently persisting.

Two ways to close it properly, and the choice is a real trade:

  • Call --configure-firewall again after Start-Service. Simple and correct, but the service's first start on a fresh install takes roughly two minutes (managed PostgreSQL bootstrap), so the installer either blocks on that or fires the second call without knowing the store is up.
  • Have the service reconcile its own firewall rule at start-up, once it has bound and therefore knows the true port. It already WARNs the exact command; doing it would need elevation the service does not have, which is precisely why the verb exists — so this probably ends as "keep warning", but it is worth deciding deliberately rather than by default.

2. A rule is opened for a surface whose store flag says disabled

PlanFirewallRules still opens a rule for an exposed surface even when the store's mcp_enabled / web_enabled is false.

That may be deliberate — a rule made ready for when the surface is enabled, so enabling it in the Viewer does not also require an elevated prompt. But it is the same shape as the stale rule #2432 just fixed: a firewall hole on a port nothing is serving. The difference is only that this one is on a port that might serve later.

Worth deciding explicitly. If it stays, the reasoning belongs in a comment where the next reader of PlanFirewallRules will find it, because the sweep logic right beside it now argues the opposite way.

Activity

  1. added a commit that references this issue on Aug 21, 2026
  2. erikdarlingdata commented on Aug 21, 2026

    @erikdarlingdata
    OwnerAuthor

    Fixed on dev by #2442, and my framing of residue 1 was wrong in the direction that matters.

    I wrote that the wrong-port rule "self-heals with one operator action, since the service's start-up check WARNs the exact command". The harm is the reverse. --configure-firewall sweeps a surface's rules for every port — SurfaceRuleWildcard rewrites PerformanceMonitor Darling MCP (port 5152) to …(port *) — before opening the current one, and the installer runs it with the service, and therefore the managed store, stopped. On an upgrade of a moved-port box, the rule that wildcard matches is the working one. So the elevated verb the operator was told to run is what closes the live LAN surface.

    That is not a one-cycle inconvenience healed by running the command; running the command is the injury. Verified in source before accepting it.

    So the sweep is now withheld when a run cannot vouch for what the surface is actually doing, and the installer re-reconciles after Start-Service on upgrades only — a fresh install's ~2-minute initdb means a second call there could only re-print the fallback about a port that is correct by definition. The comment says plainly that this narrows the window rather than closing it.

    Residue 2 is closed rather than kept, and the deciding evidence was in the product rather than in reasoning: DarlingMcpHostService's stop path already documents an admin removing the rule with --configure-firewall, and --disable-mcp sweeps it. The two verbs disagreed — --disable-mcp closed the port and the next --configure-firewall re-opened it, and every upgrade runs that verb. So a deliberately disabled endpoint had its port re-opened by an upgrade. The "ready for when it is enabled" argument I offered is written down beside the decision as rejected, with why.

    Three review rounds each caught something real, and the first is worth recording: the sweep gate initially left the Remove half open, so an --enable-mcp-enabled box — which hits Origin == File on upgrade — would have had its stale file false delete the live rule. The same class of mistake as the original defect, in the fix for it.

    One residue filed as #2445: the not-elevated branch prints a runnable command only for Open plans, so a Remove that wants its sweep hands over nothing and still returns 0. Pre-existing, but reachable more often now.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions