Repository navigation
args: remove mmap/mlock/dio flags from arg parser - #28334
Merged
Merged
Conversation
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
Member
Author
|
Server CIs are failing because the Edit: PR has been merged. |
ggerganov
approved these changes
Sep 8, 2026
ServeurpersoCom
approved these changes
Sep 8, 2026
Member
Author
|
CI failures unrelated to PR. Merging. |
1 task
zbrad
pushed a commit
to zbrad/llama.cpp
that referenced
this pull request
Sep 10, 2026
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
pl752
pushed a commit
to pl752/llama.cpp
that referenced
this pull request
Sep 15, 2026
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
2 tasks done
druide67
added a commit
to druide67/asiai-inference-server
that referenced
this pull request
Sep 16, 2026
#42) llama.cpp 0.4.1 removed --mlock/--mmap/--direct-io from its arg parser (ggml-org/llama.cpp#28334); every bundled llama.cpp daemon would refuse to start after the upgrade. --load-mode mmap+mlock is the exact equivalent and is accepted since 0.3.0. A test refuses any bundled manifest carrying a removed flag. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
zsogitbe
pushed a commit
to zsogitbe/llama.cpp
that referenced
this pull request
Sep 17, 2026
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
adderek
pushed a commit
to adderek/llama.cpp
that referenced
this pull request
Sep 28, 2026
573 upstream commits since cb30059 (2026-08-27). Conflict resolution: - qwen4exp, llama-memory-hybrid-idx: upstream merged qwen4exp itself (ggml-org#27742) and kept fixing it (ggml-org#27941 seq_cp/block keying, GDN rsqrt norm, recurrent rollback, hc ops, sparse FA). Took upstream's files and reapplied our two fork deltas: TurboQuant Q rotation on the QSA path (ce5ba4a) and the turbo -> f16 fallback for the indexer KV cache (d0e7ec2). The rest of our side was unsloth's pre-merge PR commits, superseded by upstream. - FlashAttention vec dispatch: upstream moved to ggml_cuda_get_fattn_vec_case() with per-combination GGML_CUDA_FA_<K>_<V> guards and GGML_CUDA_FA_QUANTS. TurboQuant instances stay always-compiled (appended in the CUDA and HIP CMakeLists, guards defined as 1 in fattn.cu); q8_0-q4_0 and q4_0-q8_0 added to the GGML_CUDA_FA_QUANTS default so the compiled set matches the old FA_ALL_QUANTS=OFF build. Dropped our mixed-KV restriction in the kernel selection: upstream now falls back to f16 conversion instead of aborting. Kept our RDNA3 WMMA threshold (>= 8) over upstream's new one, pending a measurement on gfx1100. - CUB: HIP keeps hipcub; MUSA follows upstream. - common/arg.cpp: upstream removed the deprecated --mmap/--mlock/--direct-io flags (ggml-org#28334); kept --hugepages. - llama-model-loader: hugepages mapping ported to the new lazy.for_file() and prefetch_size. - K2-Horizon: LLAMA_VOCAB_PRE_TYPE_K2_HORIZON renumbered to 60; n_ff_exp is now a per-layer accessor (n_ff_exp_arr). - test-llama-archs, clip.cpp: took upstream.
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
markkobo
added a commit
to markkobo/llama.cpp
that referenced
this pull request
Oct 7, 2026
…gml-org#24156) Only applies when tensors are loaded into mmap: --load-mode auto (default) and --load-mode mmap, on Linux. After CPU_REPACK copies a tensor out of mmap into its own buffer, drop the dormant file-backed source pages with madvise(MADV_DONTNEED). The mapping stays valid and re-faults safely if ever accessed. Default on because this is a correctness/efficiency fix for every default-mmap user with repack-eligible quants. Escape hatch for regressions: --no-reclaim-mmap-source. Measured on Qwen3-30B-A3B Q4_K_M (dual-socket EPYC): peak RSS 34.9 GiB -> 21.6 GiB (-13.0 GiB, -37%). Byte-identical output, identical PPL 9.0814 on wikitext-2 chunks=64. Zero major page faults. Rebased after ggml-org#28334 (load flags consolidated into --load-mode) and ggml-org#27837 (TENSOR_READ_LAZY). Plumbing updated for the new llama_model_loader ctor signature; madvise logic unchanged. Fixes ggml-org#16761
mmogr
added a commit
to mmogr/gglib
that referenced
this pull request
Oct 9, 2026
* fix(runtime): a launch with memory lock names the flag llama.cpp still has llama.cpp removed --mlock, --mmap and --direct-io in b10875 (ggml-org/llama.cpp#28334); --load-mode carries all three now. gglib passed --mlock when memory lock was on, so a launch with it on failed against any release from b10875, which today is anyone running GGLIB_LLAMA_RELEASE=latest, and would be everyone at the next pin bump. The launch passes --load-mode mmap+mlock instead. That is what --mlock meant: mmap was the default it added mlock to. --load-mode exists at the pin, b10327, so the launch is the same on both sides of the bump. The doc comments that named --mlock on the setting, the launch option and the request fields name the new flag, and the generated bindings carry the same words. * ci(llama): a script collects what a llama.cpp release changes for gglib scripts/llama_upstream.py compares the pinned release with a candidate, upstream's newest tag unless one is named, and writes what a bump would meet: the upstream commits in the subsystems gglib's behaviour depends on, llama-server --help at both tags diffed with the flags that came and went, a launch of the candidate with the flags command.rs emits on a 1 MB model up to a 200 on /health with /props kept, whether the release assets gglib downloads exist, the sampler defaults ADR 0003 defers to, and the diff of each upstream file an ADR cites, under the ADR that cites it. Every read is deterministic and bounded for an issue body; the whole of it goes to an evidence directory. The script fails loudly when command.rs emits a flag it does not know, so a new launch flag is placed before the report can vouch for it. It needs a clone of llama.cpp with its tags and no GitHub API: the newest tag is the largest b<digits> tag, and an asset's existence is a HEAD on its URL. Run against b11514 it reports --mlock gone from --help, both launches reaching /health on the new flag, every asset present and the sampler struct unchanged, which is what the fix before this commit observed by hand. * ci(llama): one issue says what moving the llama.cpp pin would meet ADR 0001 makes moving PINNED_LLAMA_RELEASE a deliberate event, and llama.cpp tags a release on nearly every merge, so the reading between two releases is a thousand commits and was not being taken: the pin sat two months and almost twelve hundred builds behind with --mlock gone from under it. llama-upstream.yml takes it every Monday, or on demand with a candidate. It runs scripts/llama_upstream.py, uploads the evidence as an artifact, and keeps one issue, titled for the pin and the newest release, whose body is replaced each run and whose title records when upstream moved. The run that finds the pin caught up closes it. It opens no pull request and touches no branch: the bump stays a commit a person makes, with the issue in front of them. The verdict, whether gglib must change and what the change is, is a reading of the evidence against gglib's code and ADRs. Claude writes it at the top of the issue when ANTHROPIC_API_KEY is set, with the files and lines each claim comes from and the ADR readings that are due; without the key the issue carries the evidence and says so. A verdict that fails to arrive is reported as missing rather than failing the run. The issue is the app's, like update-deps.yml's pull requests, and carries every label issue-labels.yml requires, since it is not a form issue for that workflow to parse. CONTRIBUTING's "Dependency updates" gains the routine, and the pin's doc comment points at it. * chore(ci): two runtime files record the lines the load-mode fix added command.rs gains the comment saying why memory lock is now spelled `--load-mode mmap+mlock` and a test that checks the flag's value as well as its presence; download/mod.rs gains the pin's pointer to the weekly upstream report. Each is still one thing, so the rows go up instead of either file splitting. --------- Co-authored-by: Matt O'Grady <172192206+mmogr@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
cont #20834
I think it has been a decent amount of time since we deprecated the
--mmap,--mlock, and--direct-ioflags in favor for--load-mode. This PR completely removes any references of the flag from the arg parser as the final cleanup of the refactor.Requirements