Skip to content

server : add lora hotswap endpoint - #8857

Merged
ngxson merged 9 commits into
ggml-org:masterfrom
ngxson:xsn/lora_server_hotswap
Aug 6, 2024
Merged

ngxson merged 9 commits into
ggml-org:masterfrom
ngxson:xsn/lora_server_hotswap

Conversation

@ngxson

@ngxson ngxson commented Aug 4, 2024 •

Copy link
Copy Markdown
Collaborator

TODO:

  • Update docs
  • Add tests

New argument: --lora-init-without-apply

If --lora-init-without-apply is specified, lora adapter will be loaded but not being apply with llama_init_from_gpt_params.

User can apply it later with the POST /lora-adapters endpoint below

New endpoints

GET /lora-adapters

Get list of all adapters. If an adapter is disabled, the scale will be set to 0.

Response:

[
    {
        "id": 0,
        "path": "my_adapter_1.gguf",
        "scale": 0.0
    },
    {
        "id": 1,
        "path": "my_adapter_2.gguf",
        "scale": 0.0
    }
]

POST /lora-adapters

Set list of adapters. To disable an adapter, either remove it from the list below, or set scale to 0.

Request:

[
  {"id": 0, "scale": 0.2},
  {"id": 1, "scale": 0.8}
]

Response:

{ "success": true }

@ngxson

ngxson commented Aug 4, 2024

Copy link
Copy Markdown
Collaborator Author

self note: maybe wait for changes from #8823 and add the list of loaded lora to struct

@Green-Sky

Copy link
Copy Markdown
Collaborator

--lora-no-apply sounds kind of contrived, maybe --lora-available or similar is better.

@ngxson

ngxson commented Aug 4, 2024 •

Copy link
Copy Markdown
Collaborator Author

I don't get what you mean. The option means "load the adapter to memory, but do not apply it right away"

probably something like --lora-apply-later or --lora-init-without-apply is more stupidly simple to understand?

@mofosyne mofosyne added the Review Complexity : Low Trivial changes to code that most beginner devs (or those who want a break) can tackle. e.g. UI fix label Aug 5, 2024
@github-actions github-actions Bot added the python python script changes label Aug 6, 2024
@ngxson
ngxson marked this pull request as ready for review August 6, 2024 11:45
@ngxson
ngxson requested a review from ggerganov August 6, 2024 11:45
@ngxson

ngxson commented Aug 6, 2024 •

Copy link
Copy Markdown
Collaborator Author

@ggerganov I added test and docs to this PR, plus adapt to change from #8823

Could you re-review this? Thank you.

@ngxson
ngxson merged commit 1e6f655 into ggml-org:master Aug 6, 2024
arthw pushed a commit to arthw/llama.cpp that referenced this pull request Aug 7, 2024
* server : add lora hotswap endpoint

* handle lora_no_apply

* fix build

* updae docs

* clean up struct def

* fix build

* add LoRA test

* fix style
@ltoniazzi ltoniazzi mentioned this pull request Aug 17, 2024
2 of 7 tasks
@ngxson ngxson changed the title server : add lora hotswap endpoint (WIP) server : add lora hotswap endpoint Aug 18, 2024
Seunghhon pushed a commit to Seunghhon/llama.cpp that referenced this pull request Apr 26, 2026
* server : add lora hotswap endpoint

* handle lora_no_apply

* fix build

* updae docs

* clean up struct def

* fix build

* add LoRA test

* fix style
ljubomirj pushed a commit to ljubomirj/llama.cpp that referenced this pull request May 6, 2026
* server : add lora hotswap endpoint

* handle lora_no_apply

* fix build

* updae docs

* clean up struct def

* fix build

* add LoRA test

* fix style
my-other-github-account pushed a commit to my-other-github-account/llama.cpp that referenced this pull request May 15, 2026
* server : add lora hotswap endpoint

* handle lora_no_apply

* fix build

* updae docs

* clean up struct def

* fix build

* add LoRA test

* fix style
AlexiAlp pushed a commit to minghaop/llama.cpp that referenced this pull request Jun 2, 2026
* server : add lora hotswap endpoint

* handle lora_no_apply

* fix build

* updae docs

* clean up struct def

* fix build

* add LoRA test

* fix style
AlexiAlp pushed a commit to minghaop/llama.cpp that referenced this pull request Jun 2, 2026
* server : add lora hotswap endpoint

* handle lora_no_apply

* fix build

* updae docs

* clean up struct def

* fix build

* add LoRA test

* fix style
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
* server : add lora hotswap endpoint

* handle lora_no_apply

* fix build

* updae docs

* clean up struct def

* fix build

* add LoRA test

* fix style
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

examples python python script changes Review Complexity : Low Trivial changes to code that most beginner devs (or those who want a break) can tackle. e.g. UI fix server

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants