Skip to content

Re-enable manual LoRA adapter free - #19983

Merged
ggerganov merged 2 commits into
ggml-org:masterfrom
PopFlamingo:fix/reenable-lora-free
Mar 18, 2026
Merged

ggerganov merged 2 commits into
ggml-org:masterfrom
PopFlamingo:fix/reenable-lora-free

Conversation

@PopFlamingo

@PopFlamingo PopFlamingo commented Feb 28, 2026 •

Copy link
Copy Markdown
Contributor

This PR proposes re-enabling manual LoRA adapter free (the llama_adapter_lora_free function), that had been previously deprecated as part of #18490.

Motivation

After reading the discussion on why llama_adapter_lora_free was deprecated and made a no-op, my understanding is that it was considered that this had no real use case, however I am currently working on a project where having the ability to unload adapters is very important due to memory constraints (mobile devices).

Without llama_adapter_lora_free, we have to fully unload and re-load the model, as well as lose any contexts (and their cached tokens) associated with it, just to free the memory associated with those LoRAs. I am not aware of any other way to achieve this.

Summary of changes

The llama_adapter_lora_free function has been re-enabled and un-deprecated; calling llama_adapter_lora_free stays optional as the LoRA will still be freed if necessary when its parent model is released. Documentation comments have been updated to reflect this new behavior.

Concretely we make sure to remove the freed LoRA from the ownership list of its owning model to prevent double frees.

Other related issues

Issue #19153 would benefit from this change as well, since I don't think the requested feature would be implementable without llama_adapter_lora_free.

@PopFlamingo
PopFlamingo requested a review from CISC as a code owner February 28, 2026 13:26
@CISC
CISC requested a review from ggerganov February 28, 2026 14:42
Comment thread include/llama.h Outdated
@PopFlamingo
PopFlamingo requested a review from ggerganov March 5, 2026 13:05
@ggerganov
ggerganov merged commit 312cf03 into ggml-org:master Mar 18, 2026
77 of 78 checks passed
Seunghhon pushed a commit to Seunghhon/llama.cpp that referenced this pull request Apr 26, 2026
* Re-enable manual LoRA adapter free

* Remove stale "all adapters must be loaded before context creation" stale comments
my-other-github-account pushed a commit to my-other-github-account/llama.cpp that referenced this pull request May 15, 2026
* Re-enable manual LoRA adapter free

* Remove stale "all adapters must be loaded before context creation" stale comments
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
* Re-enable manual LoRA adapter free

* Remove stale "all adapters must be loaded before context creation" stale comments
MrLordCat referenced this pull request in MrLordCat/llama.cpp-rdna-lab Jul 16, 2026
* Re-enable manual LoRA adapter free

* Remove stale "all adapters must be loaded before context creation" stale comments
zommiommy pushed a commit to zommiommy/llama.cpp that referenced this pull request Aug 18, 2026
* Re-enable manual LoRA adapter free

* Remove stale "all adapters must be loaded before context creation" stale comments
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* Re-enable manual LoRA adapter free

* Remove stale "all adapters must be loaded before context creation" stale comments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants