Skip to content

[Feature Request] Add RPC Backend / Remote Offload Support #181

Description

@alcoftTAO

Is your feature request related to a problem? Please describe.

Currently, llama-cpp-python exposes several compute backends (CPU, CUDA, Vulkan, SYCL, Metal), but there is no first-class support for llama.cpp's RPC backend. This makes it difficult to:

  • Offload tensor computations to a remote GPU machine.
  • Share a single GPU across multiple Python processes or containers without each one holding the full model in VRAM.
  • Run inference on machines with limited local compute (e.g., a small CPU-only node) while leveraging a powerful remote accelerator.

At the moment, users who need RPC must either build llama.cpp from source with -DGGML_RPC=ON and manage the server binary externally, or resort to fragile workarounds. A native Python API for this would be a huge quality-of-life improvement.

Describe the solution you'd like

It would be wonderful if llama-cpp-python could expose the RPC backend through its existing tensor_split / main_gpu parameters, for example:

from llama_cpp import Llama

llm = Llama(
    model_path="model.gguf",
    n_gpu_layers=-1,
    main_gpu=0,
    rpc_servers=["192.168.1.50:42232"],  # <= new kwarg
    tensor_split=[2, 1]  # <= 0 could be this machine, 1 could be the first remote server
)

Additionally, a small helper to launch the RPC server from Python would be ideal:

from llama_cpp import start_rpc_server

server = start_rpc_server(
    host="0.0.0.0",
    port=42232,
    backend="cuda",   # or "vulkan", "cpu", ...
)
# ... run inference on the client side ...
server.stop()

Even a thin wrapper around ggml-rpc-server started via subprocess would be a great start.

Describe alternatives you've considered

  • Manual build + external server: Compile llama.cpp with -DGGML_RPC=ON, run ./ggml-rpc-server & by hand, and point the client at it. Would work, but it's fragile and hard to document for new users.

Additional context

llama.cpp already has a mature RPC backend in C++ that speaks a simple TCP protocol.

Happy to help test if that would be helpful.

Activity

  1. JamePeng commented on Sep 25, 2026

    @JamePeng
    Owner

    I can try adapting the implementation to use Python as the client; since Python's server-side performance is relatively poor (which is why I prefer to discard the original author's server implementation), I favor using Python as the client to send and receive RPC requests, while launching the RPC server as a llama.cpp C++ binary for greater efficiency.

  2. alcoftTAO commented on Sep 25, 2026

    @alcoftTAO
    Author

    Looks good. Huge thanks!
    I'll be happy to help test it!

  3. JamePeng commented on Sep 29, 2026

    @JamePeng
    Owner

    I've submitted the initial implementation; feel free to give it a try.

  4. alcoftTAO commented on Sep 29, 2026

    @alcoftTAO
    Author

    I will test it in the next days! Thanks!

  5. alcoftTAO commented on Sep 30, 2026

    @alcoftTAO
    Author

    I've submitted the initial implementation; feel free to give it a try.

    Everything seems to work fine! However, I've noticed an issue when using Intel GPUs (SYCL): The generation freezes and the GPU in the remote server is not being used (but the model is loaded in VRAM).

    I'd assume it might be because of llama.cpp's RPC server (because it's still experimental), not this project.

    I'll continue testing with AMD GPUs, Vulkan, and more!

    Huge thanks! I'll close this issue now.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions