Skip to content

feat: load models and LoRA adapters from a file descriptor - #28973

Closed
ykhrustalev wants to merge 1 commit into
ggml-org:masterfrom
Liquid4All:ykhrustalev/load-from-file-descriptor
Closed

ykhrustalev wants to merge 1 commit into
ggml-org:masterfrom
Liquid4All:ykhrustalev/load-from-file-descriptor

Conversation

@ykhrustalev

@ykhrustalev ykhrustalev commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Overview

You can now load a model, or a LoRA adapter, from an already-open file descriptor
instead of a file path:

  • llama_model_load_from_fd(fd, offset, length, params)
  • llama_adapter_lora_init_from_fd(model, fd, offset, length)

The reason is Android. An app ships the GGUF inside its APK (stored with
android:noCompress), and there is no real file to open - all you get is a
descriptor for the APK plus the byte range where the model sits. Copying the
model out to disk first would double the storage use, so we read it in place.

Both functions sit on top of a new ggml call, gguf_init_from_fd(), which is a
small pread reader on top of the existing gguf_init_from_callback(). Two
things follow from that:

  • The descriptor is never closed, duplicated, or seeked. The caller keeps full
    control of it, and the same descriptor can serve more than one load at once.
  • Offsets are counted from the start of the GGUF, not from the start of the file.
    So alignment and bounds checks work the same no matter where the GGUF was
    placed inside the container.

Limits on this path: no mmap, no direct IO, no split models, and it returns
NULL on Windows.

One unrelated fix rides along: GGML_ABORT messages now also go through
GGML_LOG_ERROR, so an app sees them in its own log callback (logcat on Android)
instead of only on stderr.

Additional information

This is a port of an internal change first written against da426cb25
(2026-02-24). Upstream has changed a lot since then - it gained the
callback-based GGUF reader and the FILE * source in the model loader - so most
of the original patch was no longer needed and was dropped. What is left is about
290 lines instead of the original 900.

How it was tested:

  • tests/test-fd.cpp checks gguf_init_from_fd at offsets 0, 1, 7, 31, 32 and
    4096, rejects a length that cuts into the header or the tensor data, and
    confirms the descriptor position is untouched.
  • A second entry, test-fd-model, reuses the model that the test suite already
    downloads. It loads the same model both ways and compares the full logits, then
    does the same for a LoRA adapter. To be sure this test can actually fail, I
    removed the container offset from the read path and confirmed both cases break.
  • Ran on macOS arm64 and on a Galaxy S24 Ultra (arm64-v8a, NDK 29). 10/10 both.
  • Not built for Windows. The _WIN32 branches were only read, not compiled.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. The feature and its first implementation are mine,
    from our internal fork. An assistant did the port to current master, including
    the redesign onto gguf_init_from_callback() with pread, and wrote the tests.

Android apps ship GGUF inside the APK, where it is reachable only as an fd
plus a byte range. gguf_init_from_fd reads through pread, so the caller's fd
is never closed, duplicated or seeked, and offsets stay relative to the GGUF
start rather than to the container.

Assisted-by: Claude Opus 5
@github-actions github-actions Bot added testing Everything test related ggml changes relating to the ggml tensor library for machine learning labels Sep 16, 2026
@ykhrustalev
ykhrustalev marked this pull request as ready for review September 16, 2026 00:46
@CISC

CISC commented Sep 16, 2026

Copy link
Copy Markdown
Member

This has been discussed before, but pretty much boils down to #20402 (review)

@JohannesGaessler

Copy link
Copy Markdown
Contributor

My opinion is that we should not extend the APIs with explicit file descriptor support.

@ykhrustalev

Copy link
Copy Markdown
Contributor Author

replaced with #28993

@ykhrustalev
ykhrustalev deleted the ykhrustalev/load-from-file-descriptor branch September 17, 2026 10:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants