Repository navigation
feat: load models and LoRA adapters from a file descriptor - #28973
Closed
ykhrustalev wants to merge 1 commit into
Closed
ykhrustalev wants to merge 1 commit into
ykhrustalev wants to merge 1 commit into
Conversation
Android apps ship GGUF inside the APK, where it is reachable only as an fd plus a byte range. gguf_init_from_fd reads through pread, so the caller's fd is never closed, duplicated or seeked, and offsets stay relative to the GGUF start rather than to the container. Assisted-by: Claude Opus 5
ykhrustalev
marked this pull request as ready for review
September 16, 2026 00:46
ykhrustalev
requested review from
CISC,
JohannesGaessler and
ggerganov
as code owners
September 16, 2026 00:46
Member
|
This has been discussed before, but pretty much boils down to #20402 (review) |
Contributor
|
My opinion is that we should not extend the APIs with explicit file descriptor support. |
Contributor
Author
|
replaced with #28993 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
You can now load a model, or a LoRA adapter, from an already-open file descriptor
instead of a file path:
llama_model_load_from_fd(fd, offset, length, params)llama_adapter_lora_init_from_fd(model, fd, offset, length)The reason is Android. An app ships the GGUF inside its APK (stored with
android:noCompress), and there is no real file to open - all you get is adescriptor for the APK plus the byte range where the model sits. Copying the
model out to disk first would double the storage use, so we read it in place.
Both functions sit on top of a new ggml call,
gguf_init_from_fd(), which is asmall
preadreader on top of the existinggguf_init_from_callback(). Twothings follow from that:
control of it, and the same descriptor can serve more than one load at once.
So alignment and bounds checks work the same no matter where the GGUF was
placed inside the container.
Limits on this path: no mmap, no direct IO, no split models, and it returns
NULLon Windows.One unrelated fix rides along:
GGML_ABORTmessages now also go throughGGML_LOG_ERROR, so an app sees them in its own log callback (logcat on Android)instead of only on stderr.
Additional information
This is a port of an internal change first written against
da426cb25(2026-02-24). Upstream has changed a lot since then - it gained the
callback-based GGUF reader and the
FILE *source in the model loader - so mostof the original patch was no longer needed and was dropped. What is left is about
290 lines instead of the original 900.
How it was tested:
tests/test-fd.cppchecksgguf_init_from_fdat offsets 0, 1, 7, 31, 32 and4096, rejects a length that cuts into the header or the tensor data, and
confirms the descriptor position is untouched.
test-fd-model, reuses the model that the test suite alreadydownloads. It loads the same model both ways and compares the full logits, then
does the same for a LoRA adapter. To be sure this test can actually fail, I
removed the container offset from the read path and confirmed both cases break.
_WIN32branches were only read, not compiled.Requirements
from our internal fork. An assistant did the port to current master, including
the redesign onto
gguf_init_from_callback()withpread, and wrote the tests.