Skip to content

Process.Kill can hang indefinitely on macOS arm64 #131944

Description

@mthalman

Description

On macOS arm64, Process.Kill(entireProcessTree: true) can block indefinitely while terminating process trees created by the dotnet-watch tests.

This appeared after consuming the runtime flow containing #128598, which introduced the two-phase stop-then-kill implementation on Unix.

Enable dotnet-watch tests on mac once fixed: dotnet/sdk#55679

Reproduction Steps

  1. From a dotnet/sdk checkout at commit dc6cef6, run the dotnet-watch.Tests suite on macOS 15 arm64. In CI, this is Helix shard dotnet-watch.Tests.dll.2 on queue osx.15.arm64.open.

  2. Allow tests to dispose their child processes using Process.Kill(entireProcessTree: true).

  3. Observe that some calls never return.

A standalone runtime-only reproduction has not yet been reduced. The failure reproduces in the SDK CI shard linked below.

Expected behavior

Process.Kill(entireProcessTree: true) terminates the process tree and returns promptly.

Actual behavior

Instrumentation logged nine calls entering Process.Kill(entireProcessTree: true), but only three returned. Helix eventually detected the hang, collected dumps, and terminated the test process:

[TEST AwaitableProcess.cs:237] Killing process tree for process 18251
Hang dump timeout of '00:48:00' expired
Test application process didn't exit gracefully, exit code is '137'

One separate call returned from Kill, but its process still did not exit within 30 seconds.

Regression?

Observed after the SDK consumed the runtime flow containing #128598. That PR changed Unix process-tree termination to recursively stop the tree before killing it.

Confidence is high that the immediate hang occurs synchronously inside Process.Kill(entireProcessTree: true). Confidence that #128598 introduced the regression is moderate pending a reduced runtime-only reproduction.

Known Workarounds

None currently. Bounding the subsequent wait does not help because the call to Process.Kill itself does not return.

Configuration

  • macOS 15
  • arm64
  • Helix queue: osx.15.arm64.open
  • Work item: dotnet-watch.Tests.dll.2
  • .NET 11 runtime consumed by the SDK flow

Other information

Activity

  1. dotnet-policy-service commented on Aug 6, 2026

    @dotnet-policy-service
    Contributor

    Tagging subscribers to this area: @dotnet/area-system-diagnostics-process
    See info in area-owners.md if you want to be subscribed.

  2. adamsitnik commented on Aug 11, 2026

    @adamsitnik
    Member

    Hi @mthalman

    Thank you for a detailed bug report. I've hit hangs on macOS in #128598 but I was able to find the solution (#128598 (comment)) and get the dotnet/runtime CI green. Since then it never got stuck on macOS.

    Were you able to collect a dump locally? Or at least attach a debugger and just check which methods exactly are blocked?
    Do you happen to know if dotnet-watch.Tests.dll.2 does something unusual in terms of process management like p/invoking some native code that signs up for the SIGCHILD?

    I am asking these questions because I don't have an access to macOS arm64 machine (I will somehow try to get it, but it may take some time).

  3. added this to the 11.0.0 milestone on Aug 11, 2026
  4. mthalman commented on Aug 11, 2026

    @mthalman
    MemberAuthor

    This happens in Helix. I don't know much about Helix so I don't know if they capture a dump or not. But an example build is at https://dev.azure.com/dnceng-public/public/_build/results?buildId=1542039. The hang seems to be sporadic.

  5. added a commit that references this issue on Aug 20, 2026
  6. adamsitnik commented on Aug 20, 2026

    @adamsitnik
    Member

    My mac is useless as I can't netiher build dotnet/runtime nor run the installed dotnet executable (it's simply too old and not supported). I've tried creating a standalone repro, but I've failed to reproduce the problem: https://github.com/adamsitnik/macosrepro

    I suspect dotnet watch does something unusual that interferes with signal handling. I'll try to search for someone to try to debug it who has a working mac.

  7. slang25 commented on Oct 3, 2026

    @slang25
    Contributor

    Hi @adamsitnik, I think I've hit the same thing (or a close cousin of it), and I've got a small standalone repro that fails reliably on GitHub-hosted macOS runners, so hopefully that helps since getting hold of a mac is the hard part 🙂

    The twist on GitHub's runners is that it isn't just the call that blocks, the whole VM freezes. The runner loses contact with GitHub, step timeouts never fire, no logs get uploaded, and a root-owned process I left running as a canary stopped reporting at the same moment. Memory, network and file handles all looked healthy right up until it stopped.

    Repro

    repro.cs:

    using System.Diagnostics;
    
    for (var i = 1; i <= 200; i++)
    {
        using var p = Process.Start(new ProcessStartInfo("/bin/sh", ["-c", "sleep 60 & sleep 60; wait"]) { RedirectStandardOutput = true })!;
        Thread.Sleep(300);
        Console.WriteLine($"{DateTime.UtcNow:HH:mm:ss.fff} iteration {i}: killing tree of {p.Id}");
        p.Kill(entireProcessTree: true); // 💥
        p.WaitForExit();
    }
    Console.WriteLine("completed 200 iterations");

    Workflow:

    jobs:
      repro:
        strategy:
          matrix:
            os: [macos-26, macos-15]
        runs-on: ${{ matrix.os }}
        timeout-minutes: 4
        steps:
          - uses: actions/checkout@v5
          - uses: actions/setup-dotnet@v5
            with:
              dotnet-version: 10.0.x
          - run: dotnet run repro.cs

    Expected: 200 iterations, done in about a minute.

    Actual: the job never finishes and gets killed at the job timeout. That happened 4/4 times (2× macos-26, 2× macos-15, both arm64): https://github.com/slang25/glosharp/actions/runs/37140854273. It only takes the first few iterations, and the same loop runs all 200 iterations fine on my own Apple silicon mac (macOS 26), so it seems to be specific to the VMs.

    What I tried

    I swapped Kill(entireProcessTree: true) for other things in the same loop (https://github.com/slang25/glosharp/actions/runs/37138500101, https://github.com/slang25/glosharp/actions/runs/37139476351):

    Variant Froze
    Kill(entireProcessTree: true) 5/5
    Kill() 0/3
    just enumerating Process.GetProcesses() 0/3
    kill(SIGSTOP) then kill(SIGKILL) via P/Invoke 0/3
    kill(SIGSTOP), enumerate all processes, then kill(SIGKILL) 0/3
    find descendants from a ps -A -o pid=,ppid= snapshot, then kill(SIGKILL) each 0/3

    So it seems to need the full recursive stop/enumerate/kill, which I couldn't reproduce by hand. One thing that might be relevant to the regression label: we first saw this with our test host on .NET 8 and the repro above runs on .NET 10, so I don't think it's only #128598. I could be wrong about that though, I haven't tried a .NET 11 build.

    Images: macos-26-arm64 20260907.0351.1 (macOS 26.6.2) and macos-15 arm64.

    For now we've worked around it on macOS by killing the descendants from a ps snapshot instead: slang25/glosharp#117. Happy to run anything else on those runners if it'd help.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions