Skip to content

args: remove mmap/mlock/dio flags from arg parser - #28334

Merged
taronaeo merged 1 commit into
ggml-org:masterfrom
taronaeo:refactor/rm-dep-flags
Sep 9, 2026
Merged

taronaeo merged 1 commit into
ggml-org:masterfrom
taronaeo:refactor/rm-dep-flags

Conversation

@taronaeo

@taronaeo taronaeo commented Sep 3, 2026

Copy link
Copy Markdown
Member

Overview

cont #20834

I think it has been a decent amount of time since we deprecated the --mmap, --mlock, and --direct-io flags in favor for --load-mode. This PR completely removes any references of the flag from the arg parser as the final cleanup of the refactor.

Requirements

Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
@taronaeo
taronaeo requested review from a team, ggerganov and ngxson as code owners September 3, 2026 16:14
@taronaeo taronaeo changed the title args: officially deprecate mmap/mlock/dio flags for load mode args: remove mmap/mlock/dio flags from arg parser Sep 3, 2026
@github-actions github-actions Bot added documentation Improvements or additions to documentation server labels Sep 3, 2026
@taronaeo

taronaeo commented Sep 3, 2026 •

Copy link
Copy Markdown
Member Author

Server CIs are failing because the preset.ini file from ggml-org/test-preset-ci still uses mmap = 0. I've created a PR to change that if someone could take a look: https://huggingface.co/ggml-org/test-preset-ci/discussions/1

Edit: PR has been merged.

@taronaeo taronaeo added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 7, 2026
@taronaeo

taronaeo commented Sep 9, 2026

Copy link
Copy Markdown
Member Author

CI failures unrelated to PR. Merging.

@taronaeo
taronaeo merged commit 14a9d09 into ggml-org:master Sep 9, 2026
27 of 33 checks passed
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
druide67 added a commit to druide67/asiai-inference-server that referenced this pull request Sep 16, 2026
#42)

llama.cpp 0.4.1 removed --mlock/--mmap/--direct-io from its arg parser
(ggml-org/llama.cpp#28334); every bundled llama.cpp daemon would refuse to
start after the upgrade. --load-mode mmap+mlock is the exact equivalent and
is accepted since 0.3.0. A test refuses any bundled manifest carrying a
removed flag.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
adderek pushed a commit to adderek/llama.cpp that referenced this pull request Sep 28, 2026
573 upstream commits since cb30059 (2026-08-27). Conflict resolution:

- qwen4exp, llama-memory-hybrid-idx: upstream merged qwen4exp itself (ggml-org#27742)
  and kept fixing it (ggml-org#27941 seq_cp/block keying, GDN rsqrt norm, recurrent
  rollback, hc ops, sparse FA). Took upstream's files and reapplied our two
  fork deltas: TurboQuant Q rotation on the QSA path (ce5ba4a) and the
  turbo -> f16 fallback for the indexer KV cache (d0e7ec2). The rest of our
  side was unsloth's pre-merge PR commits, superseded by upstream.
- FlashAttention vec dispatch: upstream moved to ggml_cuda_get_fattn_vec_case()
  with per-combination GGML_CUDA_FA_<K>_<V> guards and GGML_CUDA_FA_QUANTS.
  TurboQuant instances stay always-compiled (appended in the CUDA and HIP
  CMakeLists, guards defined as 1 in fattn.cu); q8_0-q4_0 and q4_0-q8_0 added
  to the GGML_CUDA_FA_QUANTS default so the compiled set matches the old
  FA_ALL_QUANTS=OFF build. Dropped our mixed-KV restriction in the kernel
  selection: upstream now falls back to f16 conversion instead of aborting.
  Kept our RDNA3 WMMA threshold (>= 8) over upstream's new one, pending a
  measurement on gfx1100.
- CUB: HIP keeps hipcub; MUSA follows upstream.
- common/arg.cpp: upstream removed the deprecated --mmap/--mlock/--direct-io
  flags (ggml-org#28334); kept --hugepages.
- llama-model-loader: hugepages mapping ported to the new lazy.for_file()
  and prefetch_size.
- K2-Horizon: LLAMA_VOCAB_PRE_TYPE_K2_HORIZON renumbered to 60; n_ff_exp is
  now a per-layer accessor (n_ff_exp_arr).
- test-llama-archs, clip.cpp: took upstream.
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
markkobo added a commit to markkobo/llama.cpp that referenced this pull request Oct 7, 2026
…gml-org#24156)

Only applies when tensors are loaded into mmap: --load-mode auto
(default) and --load-mode mmap, on Linux.

After CPU_REPACK copies a tensor out of mmap into its own buffer, drop
the dormant file-backed source pages with madvise(MADV_DONTNEED). The
mapping stays valid and re-faults safely if ever accessed.

Default on because this is a correctness/efficiency fix for every
default-mmap user with repack-eligible quants. Escape hatch for
regressions: --no-reclaim-mmap-source.

Measured on Qwen3-30B-A3B Q4_K_M (dual-socket EPYC): peak RSS 34.9 GiB
-> 21.6 GiB (-13.0 GiB, -37%). Byte-identical output, identical PPL
9.0814 on wikitext-2 chunks=64. Zero major page faults.

Rebased after ggml-org#28334 (load flags consolidated into --load-mode) and
ggml-org#27837 (TENSOR_READ_LAZY). Plumbing updated for the new
llama_model_loader ctor signature; madvise logic unchanged.

Fixes ggml-org#16761
mmogr added a commit to mmogr/gglib that referenced this pull request Oct 9, 2026
* fix(runtime): a launch with memory lock names the flag llama.cpp still has

llama.cpp removed --mlock, --mmap and --direct-io in b10875
(ggml-org/llama.cpp#28334); --load-mode carries all three now. gglib
passed --mlock when memory lock was on, so a launch with it on failed
against any release from b10875, which today is anyone running
GGLIB_LLAMA_RELEASE=latest, and would be everyone at the next pin bump.

The launch passes --load-mode mmap+mlock instead. That is what --mlock
meant: mmap was the default it added mlock to. --load-mode exists at
the pin, b10327, so the launch is the same on both sides of the bump.

The doc comments that named --mlock on the setting, the launch option
and the request fields name the new flag, and the generated bindings
carry the same words.

* ci(llama): a script collects what a llama.cpp release changes for gglib

scripts/llama_upstream.py compares the pinned release with a candidate,
upstream's newest tag unless one is named, and writes what a bump would
meet: the upstream commits in the subsystems gglib's behaviour depends
on, llama-server --help at both tags diffed with the flags that came
and went, a launch of the candidate with the flags command.rs emits on
a 1 MB model up to a 200 on /health with /props kept, whether the
release assets gglib downloads exist, the sampler defaults ADR 0003
defers to, and the diff of each upstream file an ADR cites, under the
ADR that cites it.

Every read is deterministic and bounded for an issue body; the whole of
it goes to an evidence directory. The script fails loudly when
command.rs emits a flag it does not know, so a new launch flag is
placed before the report can vouch for it. It needs a clone of
llama.cpp with its tags and no GitHub API: the newest tag is the
largest b<digits> tag, and an asset's existence is a HEAD on its URL.

Run against b11514 it reports --mlock gone from --help, both launches
reaching /health on the new flag, every asset present and the sampler
struct unchanged, which is what the fix before this commit observed by
hand.

* ci(llama): one issue says what moving the llama.cpp pin would meet

ADR 0001 makes moving PINNED_LLAMA_RELEASE a deliberate event, and
llama.cpp tags a release on nearly every merge, so the reading between
two releases is a thousand commits and was not being taken: the pin
sat two months and almost twelve hundred builds behind with --mlock gone
from under it.

llama-upstream.yml takes it every Monday, or on demand with a
candidate. It runs scripts/llama_upstream.py, uploads the evidence as
an artifact, and keeps one issue, titled for the pin and the newest
release, whose body is replaced each run and whose title records when
upstream moved. The run that finds the pin caught up closes it. It
opens no pull request and touches no branch: the bump stays a commit a
person makes, with the issue in front of them.

The verdict, whether gglib must change and what the change is, is a
reading of the evidence against gglib's code and ADRs. Claude writes it
at the top of the issue when ANTHROPIC_API_KEY is set, with the files
and lines each claim comes from and the ADR readings that are due;
without the key the issue carries the evidence and says so. A verdict
that fails to arrive is reported as missing rather than failing the
run.

The issue is the app's, like update-deps.yml's pull requests, and
carries every label issue-labels.yml requires, since it is not a form
issue for that workflow to parse. CONTRIBUTING's "Dependency updates"
gains the routine, and the pin's doc comment points at it.

* chore(ci): two runtime files record the lines the load-mode fix added

command.rs gains the comment saying why memory lock is now spelled
`--load-mode mmap+mlock` and a test that checks the flag's value as well
as its presence; download/mod.rs gains the pin's pointer to the weekly
upstream report. Each is still one thing, so the rows go up instead of
either file splitting.

---------

Co-authored-by: Matt O'Grady <172192206+mmogr@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. server

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants