Repository navigation
server: fix router eviction races with the existing queue - #29217
Merged
ngxson merged 2 commits intoSep 22, 2026
Merged
Conversation
A model loaded by the fast path has no queue entry, so tick() evicts it at its LOADED transition before its own request is proxied. Every load now joins the queue, whose entry protects the model until its waiters leave.
A request for a model that is being stopped still sees it LOADED and is proxied into the dying child. Such a request now joins the queue and is served by the next instance. The stopping mark is cleared under the same lock that sets UNLOADED, so no request can see a model that is neither stopping nor unloaded while its child is gone.
ngxson
approved these changes
Sep 22, 2026
This was referenced Sep 23, 2026
Wizard815
pushed a commit
to Wizard815/mx-llama.cpp-Rocm10
that referenced
this pull request
Sep 29, 2026
…9217) * server: route every model load through the queue A model loaded by the fast path has no queue entry, so tick() evicts it at its LOADED transition before its own request is proxied. Every load now joins the queue, whose entry protects the model until its waiters leave. * server: do not admit requests into a stopping model A request for a model that is being stopped still sees it LOADED and is proxied into the dying child. Such a request now joins the queue and is served by the next instance. The stopping mark is cleared under the same lock that sets UNLOADED, so no request can see a model that is neither stopping nor unloaded while its child is gone. (cherry picked from commit 9919911)
LadislavSopko
pushed a commit
to 0ics-srls/llama.cpp
that referenced
this pull request
Oct 5, 2026
…9217) * server: route every model load through the queue A model loaded by the fast path has no queue entry, so tick() evicts it at its LOADED transition before its own request is proxied. Every load now joins the queue, whose entry protects the model until its waiters leave. * server: do not admit requests into a stopping model A request for a model that is being stopped still sees it LOADED and is proxied into the dying child. Such a request now joins the queue and is served by the next instance. The stopping mark is cleared under the same lock that sets UNLOADED, so no request can see a model that is neither stopping nor unloaded while its child is gone.
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
…9217) * server: route every model load through the queue A model loaded by the fast path has no queue entry, so tick() evicts it at its LOADED transition before its own request is proxied. Every load now joins the queue, whose entry protects the model until its waiters leave. * server: do not admit requests into a stopping model A request for a model that is being stopped still sees it LOADED and is proxied into the dying child. Such a request now joins the queue and is served by the next instance. The stopping mark is cleared under the same lock that sets UNLOADED, so no request can see a model that is neither stopping nor unloaded while its child is gone.
Wizard815
added a commit
to Wizard815/mx-llama.cpp-Rocm10
that referenced
this pull request
Oct 6, 2026
The conflict resolution kept the fork's ggml-org#29217 backport inside upstream's rewritten router, so the file mixed the old member set with the new declarations: server-models.cpp:1206: 'server_models::instance_t' has no member named 'th' server-models.cpp:1244: 'struct server_models' has no member named 'cv_stop' server-models.cpp:1589: no matching function for call to server_lrc_sched::pick_victim(std::unique_lock<std::mutex>&, const std::string&) server-models.cpp:1591: 'struct server_lru_sched' has no member named 'mark_slot_pending' Upstream 0.5.0 already carries ggml-org#29217 (9919911) plus the later subproc and queue refactors, and it owns the headers now: the fork changed server-models.h, server-common.h and server-http.h by zero lines, while upstream changed all three. Taking upstream's version makes the translation unit self-consistent and drops only the redundant backport. Assisted-by: Hermes Agent
edwardyoon
pushed a commit
to edwardyoon/focus-llama
that referenced
this pull request
Oct 7, 2026
…9217) * server: route every model load through the queue A model loaded by the fast path has no queue entry, so tick() evicts it at its LOADED transition before its own request is proxied. Every load now joins the queue, whose entry protects the model until its waiters leave. * server: do not admit requests into a stopping model A request for a model that is being stopped still sees it LOADED and is proxied into the dying child. Such a request now joins the queue and is served by the next instance. The stopping mark is cleared under the same lock that sets UNLOADED, so no request can see a model that is neither stopping nor unloaded while its child is gone. (cherry picked from commit 9919911)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Minimal fix for #28698 A+B repro script, one commit per variant, reusing the existing scheduler queue instead of adding a new reservation mechanism.
A) A model loaded by the fast path has no queue entry, so tick() evicts it at its LOADED transition before its own request is proxied. Every load now goes through the queue, whose entry already protects a loading model from eviction until its waiters leave.
B) A request for a model that is being stopped still sees it LOADED and is proxied into the dying child. Such a request now joins the queue and is served by the next instance, and the stopping mark is cleared under the same lock that sets UNLOADED.
The full server test suite passes locally.
Additional information
cc @ngxson, I'd like your eyes on this before we decide on the refactor.
Fixes #28698, both variants reproduced on master with the script from the issue.
Alternative to #28913.
Requirements