Skip to content

iOS 1.1.0 polls unreachable environments through Tailscale and kills the phone's tunnel every ~4 minutes #11236

Description

@Nelglor

Summary

Since updating to T3 Code iOS 1.1.0 (App Store, 2026-09-10), my iPhone's Tailscale data plane dies about 3.5 minutes after every VPN restart while T3 is connected. The trigger is T3 polling a configured environment whose tailnet peer is unreachable. Removing that environment stops the failure immediately. Adding it back with the peer unreachable reproduces it on demand.

The receive-loop death itself is a Tailscale iOS bug (tailscale/tailscale#19504, reported to Tailscale as TSS-101601), but 1.1.0 is what exposes it. 1.0.x never did on the same phone, tailnet, and offline peer.

Setup

  • iPhone, iOS 26.6.1, Tailscale 1.102.3, T3 Code 1.1.0
  • Three environments configured: bb-1 (Linux, headless server behind Tailscale Serve on :3773), GIGAPC (Linux), KatiePC (Windows)
  • Phone on home Wi-Fi, direct LAN path to bb-1 and GIGAPC
  • KatiePC with Tailscale disconnected (machine on, at the lock screen), so the peer is unreachable on the tailnet

Observed

  • T3 gets stuck on "Failed to connect. Retrying bb-1". Tailscale shows MagicSock Function Not Running: ReceiveDERP (magicsock-receive-func-error). Only a VPN off/on recovers it.
  • Tailscale's own log on the phone, in the minute before each failure, shows T3 opening new TCP connections to KatiePC's t3 port roughly every 25 seconds (Accept: TCP{phone:port > <katiepc-tailnet-ip>:443}, then open-conn-track: timeout opening ... to node [<katiepc>]), each one driving WireGuard handshake retries to a peer the DERP relay doesn't know.
  • Server-side, bb-1 sees the phone fall back to relay-only discovery pings and never answer direct pings until the phone's Tailscale restarts.

Evidence

  • Week before the update: 0 phone VPN restarts per day, phone equally active on the tailnet each day.
  • Day of the update: 24 restarts, 30 stall bursts, all inside T3 sessions.
  • Same day, 7 hours with the phone on the tailnet and T3 closed: 0 stalls.
  • Restart to next stall: median 214 s, typical 180 to 260 s (n=26).
  • Removed KatiePC from T3 while it was offline: 0 stalls in 28 minutes of active use (baseline one every 4 minutes), and none overnight.
  • Re-added KatiePC with its t3 server stopped but Tailscale still up on it: clean. Then disconnected Tailscale on KatiePC: stall 3.5 minutes later.

So the trigger is "environment whose tailnet peer is unreachable", not "environment whose server is down".

Ask

Happy to test a TestFlight build; reproduction takes about four minutes.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions