Skip to content

server: fix router eviction races with the existing queue - #29217

Merged
ngxson merged 2 commits into
ggml-org:masterfrom
ServeurpersoCom:server-router-eviction-race
Sep 22, 2026
Merged

ngxson merged 2 commits into
ggml-org:masterfrom
ServeurpersoCom:server-router-eviction-race

Conversation

@ServeurpersoCom

@ServeurpersoCom ServeurpersoCom commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Minimal fix for #28698 A+B repro script, one commit per variant, reusing the existing scheduler queue instead of adding a new reservation mechanism.

A) A model loaded by the fast path has no queue entry, so tick() evicts it at its LOADED transition before its own request is proxied. Every load now goes through the queue, whose entry already protects a loading model from eviction until its waiters leave.

B) A request for a model that is being stopped still sees it LOADED and is proxied into the dying child. Such a request now joins the queue and is served by the next instance, and the stopping mark is cleared under the same lock that sets UNLOADED.

The full server test suite passes locally.

Additional information

cc @ngxson, I'd like your eyes on this before we decide on the refactor.

Fixes #28698, both variants reproduced on master with the script from the issue.

Alternative to #28913.

Requirements

A model loaded by the fast path has no queue entry, so tick() evicts
it at its LOADED transition before its own request is proxied. Every
load now joins the queue, whose entry protects the model until its
waiters leave.
A request for a model that is being stopped still sees it LOADED and
is proxied into the dying child. Such a request now joins the queue
and is served by the next instance. The stopping mark is cleared
under the same lock that sets UNLOADED, so no request can see a
model that is neither stopping nor unloaded while its child is gone.
@ngxson
ngxson merged commit 9919911 into ggml-org:master Sep 22, 2026
12 checks passed
Wizard815 pushed a commit to Wizard815/mx-llama.cpp-Rocm10 that referenced this pull request Sep 29, 2026
…9217)

* server: route every model load through the queue

A model loaded by the fast path has no queue entry, so tick() evicts
it at its LOADED transition before its own request is proxied. Every
load now joins the queue, whose entry protects the model until its
waiters leave.

* server: do not admit requests into a stopping model

A request for a model that is being stopped still sees it LOADED and
is proxied into the dying child. Such a request now joins the queue
and is served by the next instance. The stopping mark is cleared
under the same lock that sets UNLOADED, so no request can see a
model that is neither stopping nor unloaded while its child is gone.

(cherry picked from commit 9919911)
LadislavSopko pushed a commit to 0ics-srls/llama.cpp that referenced this pull request Oct 5, 2026
…9217)

* server: route every model load through the queue

A model loaded by the fast path has no queue entry, so tick() evicts
it at its LOADED transition before its own request is proxied. Every
load now joins the queue, whose entry protects the model until its
waiters leave.

* server: do not admit requests into a stopping model

A request for a model that is being stopped still sees it LOADED and
is proxied into the dying child. Such a request now joins the queue
and is served by the next instance. The stopping mark is cleared
under the same lock that sets UNLOADED, so no request can see a
model that is neither stopping nor unloaded while its child is gone.
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…9217)

* server: route every model load through the queue

A model loaded by the fast path has no queue entry, so tick() evicts
it at its LOADED transition before its own request is proxied. Every
load now joins the queue, whose entry protects the model until its
waiters leave.

* server: do not admit requests into a stopping model

A request for a model that is being stopped still sees it LOADED and
is proxied into the dying child. Such a request now joins the queue
and is served by the next instance. The stopping mark is cleared
under the same lock that sets UNLOADED, so no request can see a
model that is neither stopping nor unloaded while its child is gone.
Wizard815 added a commit to Wizard815/mx-llama.cpp-Rocm10 that referenced this pull request Oct 6, 2026
The conflict resolution kept the fork's ggml-org#29217 backport inside upstream's
rewritten router, so the file mixed the old member set with the new
declarations:

  server-models.cpp:1206: 'server_models::instance_t' has no member named 'th'
  server-models.cpp:1244: 'struct server_models' has no member named 'cv_stop'
  server-models.cpp:1589: no matching function for call to
      server_lrc_sched::pick_victim(std::unique_lock<std::mutex>&, const std::string&)
  server-models.cpp:1591: 'struct server_lru_sched' has no member named 'mark_slot_pending'

Upstream 0.5.0 already carries ggml-org#29217 (9919911) plus the later subproc and
queue refactors, and it owns the headers now: the fork changed
server-models.h, server-common.h and server-http.h by zero lines, while
upstream changed all three. Taking upstream's version makes the translation
unit self-consistent and drops only the redundant backport.

Assisted-by: Hermes Agent
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
…9217)

* server: route every model load through the queue

A model loaded by the fast path has no queue entry, so tick() evicts
it at its LOADED transition before its own request is proxied. Every
load now joins the queue, whose entry protects the model until its
waiters leave.

* server: do not admit requests into a stopping model

A request for a model that is being stopped still sees it LOADED and
is proxied into the dying child. Such a request now joins the queue
and is served by the next instance. The stopping mark is cleared
under the same lock that sets UNLOADED, so no request can see a
model that is neither stopping nor unloaded while its child is gone.

(cherry picked from commit 9919911)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Eval bug: tick() evicts models which are being requested -> requests fail due to http client error

2 participants