Repository navigation
gguf : align the data section relative to the GGUF start, not the file - #28993
Merged
CISC merged 5 commits intoSep 17, 2026
Merged
Conversation
gguf_init_from_file_ptr reads a GGUF from the current file position, but padded the data section from file offset 0, so a GGUF embedded at an offset that is not a multiple of the alignment loaded without error and returned wrong tensor data. Also adds llama_adapter_lora_init_from_file_ptr, and disables mmap with a warning when an embedded data section is not aligned, instead of asserting in ggml. Assisted-by: Claude Opus 5
The test now checks that mmap is disabled only for an unaligned offset. Assisted-by: Claude Fable 5.1
ykhrustalev
marked this pull request as ready for review
September 16, 2026 14:52
ykhrustalev
requested review from
CISC,
JohannesGaessler and
ggerganov
as code owners
September 16, 2026 14:52
Comment on lines
+688
to
+694
| // mmap places tensors at their file offsets, so an embedded GGUF must be aligned in the file too | ||
| const size_t tensor_align = ggml_backend_buft_get_alignment(ggml_backend_cpu_buffer_type()); | ||
| if (use_mmap && gguf_get_data_offset(metadata) % tensor_align != 0) { | ||
| LLAMA_LOG_WARN("%s: GGUF data section at file offset %zu is not %zu byte aligned, mmap is disabled\n", | ||
| __func__, gguf_get_data_offset(metadata), tensor_align); | ||
| use_mmap = false; | ||
| } |
Contributor
There was a problem hiding this comment.
This should throw a runtime error instead.
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
JohannesGaessler
approved these changes
Sep 16, 2026
CISC
approved these changes
Sep 17, 2026
CISC
pushed a commit
that referenced
this pull request
Sep 17, 2026
#28993) * gguf : align the data section relative to the GGUF start, not the file gguf_init_from_file_ptr reads a GGUF from the current file position, but padded the data section from file offset 0, so a GGUF embedded at an offset that is not a multiple of the alignment loaded without error and returned wrong tensor data. Also adds llama_adapter_lora_init_from_file_ptr, and disables mmap with a warning when an embedded data section is not aligned, instead of asserting in ggml. Assisted-by: Claude Opus 5 * llama : load lora from path through the FILE* variant The test now checks that mmap is disabled only for an unaligned offset. Assisted-by: Claude Fable 5.1 * Update ggml/src/gguf.cpp Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Update include/llama.h Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * llama : error on unaligned mmap of an embedded GGUF, drop test-load-file-ptr --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
adromir
pushed a commit
to adromir/llama-cpp-turboquant
that referenced
this pull request
Sep 17, 2026
ggml-org#28993) * gguf : align the data section relative to the GGUF start, not the file gguf_init_from_file_ptr reads a GGUF from the current file position, but padded the data section from file offset 0, so a GGUF embedded at an offset that is not a multiple of the alignment loaded without error and returned wrong tensor data. Also adds llama_adapter_lora_init_from_file_ptr, and disables mmap with a warning when an embedded data section is not aligned, instead of asserting in ggml. Assisted-by: Claude Opus 5 * llama : load lora from path through the FILE* variant The test now checks that mmap is disabled only for an unaligned offset. Assisted-by: Claude Fable 5.1 * Update ggml/src/gguf.cpp Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Update include/llama.h Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * llama : error on unaligned mmap of an embedded GGUF, drop test-load-file-ptr --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
Te-eMster
pushed a commit
to Te-eMster/mx-llama.cpp
that referenced
this pull request
Sep 18, 2026
ggml-org#28993) * gguf : align the data section relative to the GGUF start, not the file gguf_init_from_file_ptr reads a GGUF from the current file position, but padded the data section from file offset 0, so a GGUF embedded at an offset that is not a multiple of the alignment loaded without error and returned wrong tensor data. Also adds llama_adapter_lora_init_from_file_ptr, and disables mmap with a warning when an embedded data section is not aligned, instead of asserting in ggml. Assisted-by: Claude Opus 5 * llama : load lora from path through the FILE* variant The test now checks that mmap is disabled only for an unaligned offset. Assisted-by: Claude Fable 5.1 * Update ggml/src/gguf.cpp Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Update include/llama.h Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * llama : error on unaligned mmap of an embedded GGUF, drop test-load-file-ptr --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
fencerJP
pushed a commit
to fencerJP/llama-apu
that referenced
this pull request
Sep 19, 2026
ggml-org#28993) * gguf : align the data section relative to the GGUF start, not the file gguf_init_from_file_ptr reads a GGUF from the current file position, but padded the data section from file offset 0, so a GGUF embedded at an offset that is not a multiple of the alignment loaded without error and returned wrong tensor data. Also adds llama_adapter_lora_init_from_file_ptr, and disables mmap with a warning when an embedded data section is not aligned, instead of asserting in ggml. Assisted-by: Claude Opus 5 * llama : load lora from path through the FILE* variant The test now checks that mmap is disabled only for an unaligned offset. Assisted-by: Claude Fable 5.1 * Update ggml/src/gguf.cpp Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Update include/llama.h Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * llama : error on unaligned mmap of an embedded GGUF, drop test-load-file-ptr --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
This was referenced Sep 23, 2026
1 task done
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
ggml-org#28993) * gguf : align the data section relative to the GGUF start, not the file gguf_init_from_file_ptr reads a GGUF from the current file position, but padded the data section from file offset 0, so a GGUF embedded at an offset that is not a multiple of the alignment loaded without error and returned wrong tensor data. Also adds llama_adapter_lora_init_from_file_ptr, and disables mmap with a warning when an embedded data section is not aligned, instead of asserting in ggml. Assisted-by: Claude Opus 5 * llama : load lora from path through the FILE* variant The test now checks that mmap is disabled only for an unaligned offset. Assisted-by: Claude Fable 5.1 * Update ggml/src/gguf.cpp Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Update include/llama.h Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * llama : error on unaligned mmap of an embedded GGUF, drop test-load-file-ptr --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Rework of #28973 using
FILE *instead of file descriptors, as asked there and in #20402.#28973 can be closed.
gguf_init_from_file_ptr()reads from the current file position, but aligns the datasection from offset 0. So a GGUF stored at an unaligned offset inside a bigger file (an
Android APK asset, for example) loads without error and returns wrong tensor data. The
reader now aligns from where the GGUF starts.
With that fixed, no new API is needed for embedded models:
fseekto the GGUF, thenllama_model_load_from_file_ptr(). On Android:Two small additions:
llama_adapter_lora_init_from_file_ptr(), since LoRA had noFILE *entry point.instead of ggml asserting.
Additional information
test-gguf: newfile_offsetmode, GGUF written after 7 junk bytes. Fails withoutthe fix.
test-load-file-ptr(needs the model fixture): model and LoRA loaded from an offset,logits compared with a normal load. It is a new test file; drop it if not wanted.
LoRA asset landed at offset mod 32 = 16.
Requirements
run the checks.