Skip to content

Fix ggml_metal crash on app exit (llama.cpp Metal cleanup) #1063

Description

@malibio

Problem

The app crashes on exit with a ggml_metal_device_init / ggml_metal_device_free backtrace from llama.cpp's Metal backend cleanup. This happens consistently during app shutdown:

WARNING: GGML_BACKTRACE_LLDB may cause native MacOS Terminal.app to crash.
0   nodespace-app    ggml_print_backtrace + 276
1   nodespace-app    ggml_abort + 156
2   nodespace-app    ggml_metal_device_init + 0
3   nodespace-app    ggml_metal_device_free + 24
4   nodespace-app    vector<unique_ptr<ggml_metal_device>>~D1Ev + 72
5   libsystem_c.dylib  __cxa_finalize_ranges + 480
6   libsystem_c.dylib  exit + 44

The crash occurs in the C++ destructor chain during process exit (__cxa_finalize_ranges), when llama.cpp's static Metal device vector is destroyed. The Metal context has already been torn down by the time the static destructor runs, causing a use-after-free abort.

Impact

  • Crash report dialog appears on every app close
  • No data loss (crash happens after all Rust cleanup completes)
  • Bad UX — users see a crash every time they quit the app

Likely Root Cause

llama.cpp stores Metal devices in a static vector<unique_ptr<ggml_metal_device>>. During process exit, C++ static destructors run in reverse order. If the Metal framework has already been cleaned up by macOS before the static vector destructor runs, ggml_metal_device_free accesses freed Metal resources and aborts.

This is a known pattern in llama.cpp: ggml-org/llama.cpp#17869

Proposed Solution

  1. Explicitly drop llama.cpp contexts before process exit — ensure all LlamaModel, LlamaContext, and related objects are dropped in our Rust shutdown sequence, before the process exits and C++ static destructors run
  2. Check if our llama.cpp binding version has a fix — the linked PR may address this
  3. As a fallback: suppress the crash dialog on macOS by catching the signal during shutdown

Activity

  1. malibio commented on Apr 7, 2026

    @malibio
    CollaboratorAuthor

    Baseline: 128 test files passed, 3918 tests passed

  2. self-assigned this
    on Apr 7, 2026
  3. added 4 commits that reference this issue on Apr 7, 2026
    d735db8
    6226e58
    3e5adbb
    4a8f9fe
  4. malibio commented on Apr 7, 2026

    @malibio
    CollaboratorAuthor

    ✅ Issue #1063 resolved and merged to main via PR #1065

    Summary

    Fixed the ggml_metal crash on app exit by implementing proper cleanup ordering:

    1. Reset ManagedAgentState inference engine to NoOp (drops ChatEngine Arc references)
    2. Unload chat model via GgufModelManager
    3. Release embedding GPU resources
    4. Release global llama backend

    Verification

    The fix ensures ChatEngine Arc is explicitly dropped before Metal backend teardown, preventing use-after-free crashes during shutdown.

  5. added 3 commits that reference this issue on Jun 22, 2026
    feee7a7
    c1b3a73
    6539486
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

backendbugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions