Skip to content

disagg: RPC batching, KV prefetch streaming, sync fence, disagg fixes - #5

Open
b-ostrov wants to merge 13 commits into
ServeurpersoCom:poc/server-disagg-prefill-decodefrom
b-ostrov:disagg-rpc-improvements
Open

b-ostrov wants to merge 13 commits into
ServeurpersoCom:poc/server-disagg-prefill-decodefrom
b-ostrov:disagg-rpc-improvements

Conversation

@b-ostrov

@b-ostrov b-ostrov commented Jul 8, 2026

Copy link
Copy Markdown

Hi, thanks for the disagg prefill/decode work — I've been running it on a direct 10GbE link between a Mac Studio (M2 Ultra, Metal decode) and a DGX Spark (GB10, CUDA prefill), on Qwen3.6-27B and Qwen3.6-35B-A3B. This PR is a set of patches on top of poc/server-disagg-prefill-decode (based on f4a0744), plus the numbers I measured along the way, in case any of it is useful upstream.

Patches

  1. has_mtmd = !files.empty() in process_mtmd_prompt (server-common.cpp) — as written, having an --mmproj loaded unconditionally disabled disagg prefill even for pure-text requests. This mattered a lot since vision models are common now (Qwen3.6-VL etc.) and I didn't want to give up disagg just because the model can see images.

  2. Skip disagg when the slot has a cached common prefix (server-context.cpp) — disagg re-establishes the whole prefix KV on every request by design, which is right for a cold prompt but wasteful for multi-turn conversations where most of the prefix is already resident locally. Added a check on get_common_prefix() before triggering the disagg path, so follow-up turns extend the local KV instead of re-shipping everything through the remote prefill device.

  3. Batched RPC tensor fetch — new RPC_CMD_GET_TENSOR_BATCH (proto v4.1). The KV handoff after disagg prefill was doing one round-trip per tensor, which capped wire throughput at ~3.8 Gbps on our 10GbE link regardless of MTU/buffer tuning. Batching the fetch into one round trip got us to ~9.8 Gbps (essentially wire speed).

  4. 32MB SO_RCVBUF/SNDBUF on the RPC sockets (transport.cpp) — default OS buffers were bottlenecking the bulk tensor transfers independent of (3).

  5. ggml_backend_rpc_synchronize is currently a no-op — graph_compute over RPC is fire-and-forget (no response expected), so there's nothing to block on. I think this is worth flagging as a correctness issue, not just a perf one: on a slot doing multiple RPC graphs back to back (e.g. disagg prefill chunks), a caller relying on synchronize as a barrier can read a buffer before the queued work has actually completed on the remote side. I made it issue a cheap round-trip command (DEVICE_COUNT) as a real fence — since the remote processes commands in order, getting a response means the queue up to that point has drained.

  6. KV prefetch streaming during prefill — this was the biggest win. Added llama_state_seq_prefetch_ext → llama_memory_i::state_prefetch (implemented for llama_kv_cache, forwarded through llama_memory_hybrid — Qwen3.6 is a hybrid recurrent/attention architecture, so the forwarding turned out to be required, not optional) → ggml_backend_rpc_prefetch_tensor_batch, a fire-and-forget prefetch request whose response is drained lazily into a client-side region cache keyed by (remote tensor ptr, offset). Called per-chunk inside disagg_prefill_batch, with the per-chunk llama_synchronize removed since queue ordering is enough to guarantee correctness. Net effect: the attention-KV transfer for a chunk is fully hidden behind the next chunk's remote compute, instead of happening as a separate blocking step after prefill finishes.

Numbers (Qwen3.6-27B UD-Q4_K_XL, Mac Studio M2 Ultra ↔ DGX Spark GB10, direct 10GbE, MTU 8000)

KV handoff, 28.8K-token prompt: 1264ms → 544ms after patch 6 (86% of the ~1.1GB transfer served from the prefetch cache; the residual ~150ms is the fixed 149.6MiB recurrent DeltaNet state, which is not prefetchable, plus the final chunk's compute).

End-to-end, same 28.8K-token prompt: prefill 48.9s on Spark (~590 tok/s) + 0.57s handoff + 31.3s decode on Mac (~12.8 tok/s at that depth).

Cluster vs. single-node (same ~2200-token prompt, temp=0, same model):

Mode Prompt processing Decode Total (300 tok)
Cluster (Spark prefill + Mac decode) 4.4s 26.8 tok/s 15.6s
Mac only (Metal, no RPC) 8.3s 24.2 tok/s 20.7s
Spark only (CUDA via RPC, --device RPC0) 4.3s 9.1 tok/s 37.3s

Spark-only decode being ~2.7x slower than Mac isn't a disagg issue — it's GB10 being memory-bandwidth-bound on single-batch autoregressive decode, consistent with what others have reported for this hardware. The disaggregated split plays to each device's strength: CUDA for parallel prefill, Metal for decode.

Depth scaling (llama-bench, tg128, cluster prefill device):

Depth pp512 (tok/s) tg128 (tok/s)
0 333-335 24.9-25.0
16384 246-250 16.0-16.2
65536 133-134 7.8

Speculative decoding (--spec-type ngram-mod) stacks fine with disagg, but with a caveat: on the first disagg request in a session, draft-accept rate is low (observed ~9%) because the CUDA prefill and Metal decode produce numerically slightly different logits at temp=0 (verified — outputs diverge starting a few dozen tokens in on identical greedy-decode requests), so n-grams learned from one don't transfer cleanly to the other on that first pass. On follow-up turns disagg is skipped anyway (see patch 2), so ngram-mod runs at full effectiveness there. Net gain on repetitive/code-edit workloads: +39% (24→33.4 tok/s) on a code-edit task, up to 2.1x on highly repetitive prompts, neutral (no regression) on fresh text.

Layer-split mode (--device MTL0,RPC0 --split-mode layer --tensor-split, no --prefill-device) for models too large for either machine alone — tested on Qwen3.6-35B-A3B-Q8 (34GB) split 55/45: decode 18.6 tok/s vs. 65.6 tok/s running the same quant Mac-only. Expected — every token's activations now cross the 10GbE link — but it's the only way to run something that genuinely doesn't fit in 147GB (Mac) or 119GB (Spark) alone, giving ~260GB combined. Not something I'd use unless the model actually requires it.

Network / deployment notes (not code changes, but cost me real debugging time)

  • Direct link, no switch: two 10GbE NICs cabled point-to-point on a dedicated /24, separate from the management LAN used for SSH. Keeps the RPC path off any shared network.
  • MTU 8000, not 9000: the Spark's onboard Realtek NIC caps out around 8178 bytes — MTU 9000 silently doesn't pass end-to-end on that chip. 8000 is the largest value that reliably works on both ends.
  • Socket tuning beyond the buffer patch above: net.core.wmem_max/rmem_max = 16MB and net.ipv4.tcp_slow_start_after_idle=0 on the Spark side (persisted via /etc/sysctl.d/), kern.ipc.maxsockbuf=16777216 on the Mac (/etc/sysctl.conf) plus networksetup -setMTU. Without tcp_slow_start_after_idle=0 specifically, throughput on a request after any idle gap re-ramps from a slow-start window instead of going straight to line rate — noticeable since our traffic pattern is bursty (idle between requests, then a burst per prefill).
  • macOS "Local Network" permission silently breaks RPC — worth a README callout for other Mac users of this fork. On macOS 15.x, any non-system process (including a plain llama-server binary run as a normal user) that tries to connect() to the RPC port gets errno 65 ("No route to host") even though the host is reachable — nc/system tools connect fine, confirming it's an app-level Local Network permission gate, not routing. There's no user-facing prompt for this in our testing; it just fails. The workaround we're running with is launching llama-server from a root-owned LaunchDaemon (root is exempt from the check), but that's a workaround, not a fix — a note in the README about this would save the next Mac user a confusing debugging session, since the error message points at routing/firewall, not permissions.
  • Both binaries must be rebuilt together on any RPC protocol change — the hello handshake checks major equal / server minor <= client, so a stale rpc-server on the remote box fails closed rather than degrading, which is good, but easy to trip over when only rebuilding one side.
  • Reliability: rpc-server on the CUDA box under a systemd unit (Restart=always) and the router llama-server on macOS under the LaunchDaemon mentioned above — both survive a reboot and a crash without manual intervention. Worth doing before treating this as more than a demo, since a single stuck RPC connection previously required killing and restarting both processes.

Scope of this PR

Self-contained on top of f4a0744, ~15 files touched across ggml/src/ggml-rpc/, src/llama-kv-cache.{cpp,h}, src/llama-memory-hybrid.{cpp,h}, and tools/server/server-context.cpp. Happy to split it up by concern if you'd rather review RPC batching / prefetch streaming / disagg fixes separately — they're fairly independent of each other.

ServeurpersoCom and others added 11 commits June 22, 2026 06:08
Spread the disaggregated prefill over update_slots ticks instead of
blocking the loop. Add SLOT_STATE_DISAGG_PREFILL and a per slot cursor,
split disagg_prefill into begin/step/finish, and step one chunk per tick
so other slots decode on ctx_tgt between chunks. The handoff and the
last token run on return via the existing prompt tail.

Reset n_empty_consecutive on prefill progress so the empty batch
watchdog does not abort while a lone slot prefills on ctx_pfx.

Output is token for token identical to local decode.
Replace the single fixed prefill sequence with a pool of n_prefill
sequences on ctx_pfx. A slot acquires a free sequence when it enters
disaggregated prefill and releases it on handoff, so up to n_prefill
prefills run concurrently, each on its own sequence. When the pool is
full the slot falls back to local prompt processing.

Move the per chunk trace to SLT_DBG and drop its timing. The per
request summary stays at SLT_INF.
The disaggregated prefill path does not hand off the draft context KV, so
it is skipped when a draft context is active. Speculative decoding and
disagg prefill are mutually exclusive for now.

Document --prefill-device, --n-prefill and --n-prefill-chunk in the server
README, with a disaggregated prefill/decode section and an RPC example.
The disagg prefill path went through the normal prompt tail with n_past = 0,
which dropped the prefix and erased the KV just handed off into ctx_tgt, so
the prompt was reprocessed locally and the remote prefill was wasted.

Set n_past to the prefilled prefix length and skip the checkpoint/pos_min
reset for disagg, so ctx_tgt only evaluates the last token. Split the slot
flag into disagg_attempted (try once, no retry loop) and disagg_prefilled
(handoff valid), set the latter only after the state get/set succeed, check
their returned sizes and fall back to local processing on failure. Clamp
n-prefill-chunk to at least 1.
disagg_prefill_batch packs one chunk from every prefilling slot into a
single llama_decode on ctx_pfx, replacing the per slot disagg_prefill_step
that issued one decode per slot. it runs once before the prompt batching
loop, and slots that complete their prefix flip to SLOT_STATE_STARTED so
the KV handoff and the final token run on ctx_tgt the same tick.

a failed ctx_pfx decode falls every batched slot back to local prompt
processing: disagg_pos stays put so disagg_prefill_finish releases the
sequence without a handoff and disagg_prefilled stays false.
gate disagg prefill on SERVER_TASK_TYPE_COMPLETION so embedding and rerank
keep their full ctx_tgt outputs, and skip it when the slot carries lora
adapters since ctx_pfx never sets them and a lora prefix KV would be wrong.

in disagg_prefill_batch only the slots packed into the ctx_pfx batch
advance and fall back on a failed decode, guarded against an empty batch.
create ctx_pfx with the same n_seq_max as ctx_tgt so both contexts share
the same n_stream, otherwise llama_state_seq_set_data rejects the handoff
with an n_stream mismatch and every prefill silently falls back to local.

clamp n_prefill to n_parallel since the pool cannot hand off more
concurrent prefills than there are decode slots.
Six independent patches on top of the disaggregated prefill/decode
server, developed and benchmarked running Qwen3.6-27B/35B-A3B across a
direct 10GbE link (Mac Studio M2 Ultra decode <-> DGX Spark GB10 CUDA
prefill):

- process_mtmd_prompt: has_mtmd = !files.empty(), so a loaded --mmproj
  no longer unconditionally disables disagg prefill for text-only
  requests.
- server-context: skip the disagg path when the slot already has a
  cached common prefix, so multi-turn conversations extend the local
  KV instead of re-shipping the whole prefix through the remote
  prefill device on every turn.
- New RPC_CMD_GET_TENSOR_BATCH (proto v4.1): batches the post-prefill
  KV fetch into one round trip instead of one per tensor. Raised
  measured wire throughput on our 10GbE link from ~3.8 to ~9.8 Gbps.
- transport: 32MB SO_RCVBUF/SNDBUF on RPC sockets, independent win on
  top of the batching above.
- ggml_backend_rpc_synchronize was a no-op (graph_compute over RPC is
  fire-and-forget). Made it issue a cheap round-trip command as a real
  fence -- flagging this as a correctness fix, not just perf: a caller
  chaining multiple RPC graphs (e.g. disagg prefill chunks) could
  previously read a buffer before queued remote work had completed.
- KV prefetch streaming during prefill: llama_state_seq_prefetch_ext
  -> llama_memory_i::state_prefetch (implemented for llama_kv_cache,
  forwarded through llama_memory_hybrid -- required for Qwen3.6's
  hybrid recurrent/attention architecture) -> a fire-and-forget
  ggml_backend_rpc_prefetch_tensor_batch request, drained lazily into
  a client-side region cache keyed by (remote tensor ptr, offset).
  Hides the attention-KV transfer for each prefill chunk behind the
  next chunk's remote compute. Measured: 28.8K-token prompt KV handoff
  1264ms -> 544ms (86% served from prefetch cache).

See PR description for full benchmark tables (cluster vs. single-node,
depth scaling, speculative decoding interaction, layer-split mode) and
deployment notes.
denmrnngp-cloud added 2 commits July 10, 2026 12:22
…I config dialog with slider

Assisted-by: Claude
…ements

# Conflicts:
#	tools/server/server-context.cpp
#	tools/server/server-models.cpp
#	tools/server/server-models.h
#	tools/ui/src/lib/components/app/models/ModelsSelectorOption.svelte
#	tools/ui/src/lib/constants/api-endpoints.ts
#	tools/ui/src/lib/services/models.service.ts
@ServeurpersoCom

Copy link
Copy Markdown
Owner

Interesting; I'm going to rebase my branch to see more !

@ServeurpersoCom
ServeurpersoCom force-pushed the poc/server-disagg-prefill-decode branch from f4a0744 to 429802c Compare July 13, 2026 10:16
@ServeurpersoCom

Copy link
Copy Markdown
Owner

I've rebased on my end. I'm going to rebase yours to see.

@ServeurpersoCom

Copy link
Copy Markdown
Owner

OK the prefetch path looks like a hack to me, we need real async primitives in libllama for this. I rebased your commit here: https://github.com/ServeurpersoCom/llama.cpp/tree/disagg-rpc-improvements / ggml-org@35dc99d
And confirm numbers were measured on real hardware.

@b-ostrov

Copy link
Copy Markdown
Author

Yes, all numbers are from real hardware: Mac Studio M2 Ultra (192GB) + DGX Spark GB10, direct 10GbE link, MTU 8000. Happy to re-run any specific benchmark if useful

@ServeurpersoCom

Copy link
Copy Markdown
Owner

The official discussion can be found here : ggml-org#21266
You can look/try this one : https://github.com/ServeurpersoCom/llama.cpp/tree/poc/server-disagg-prefill-decode2

@b-ostrov

Copy link
Copy Markdown
Author

ran your decode2 on the same hardware - e2e parity with my branch (45.5s prompt / 13.3 tok/s decode on a 26.5K prompt), but your get(net) is 1.3s vs 0.5s on mine - the 32MB socket buffers alone should close most of that gap

@ServeurpersoCom

ServeurpersoCom commented Jul 14, 2026 •

Copy link
Copy Markdown
Owner

Look this one : ggml-org#25675
No prefill streaming but I am starting to wonder if streaming is not plain overengineering here, and whether a few seconds of transfer per request over a 1Gbps LAN is simply acceptable in practice...
Need more testing !

We shouldn't hesitate to test all possible implementations in order to extract the best approach for llama.cpp even if it means combining different ideas!

Also, you really shouldn’t trust agentic benchmarks without asking for a simple, reproducible test script and running it yourself. Things can quickly go off the rails: the cobra effect is always hidden and full of surprises.

@b-ostrov

b-ostrov commented Jul 14, 2026 •

Copy link
Copy Markdown
Author

@ServeurpersoCom
@felix314159

tested all three on the same rig (M2 Ultra <> GB10, 10GbE, 26.5K prompt): #25675 blocking = 64-65s e2e, my streaming = 68.1s, decode2 = 68.5s. On a fast link blocking actually wins - the dedicated prefill worker beats chunked scheduling by more than streaming saves. On 1Gbps the math flips: 1GB KV = ~10s serialized per request.

Suggests: #25675's worker design + streaming as an option for slow links + the multi-turn cache skip

@ServeurpersoCom

Copy link
Copy Markdown
Owner

It's still better to work on upstream (ggml-org) than on my fork!

@b-ostrov

b-ostrov commented Jul 14, 2026 •

Copy link
Copy Markdown
Author

Well, it was a good try

What the hybrid was:

Base: this PR (ggml-org#25675) - the dedicated prefill worker design

  • async RPC transport from poc/server-disagg-prefill-decode2 - cherry-picked the get_tensor_async implementation over GET_TENSOR, the real synchronize fence, and the deferred-read path in llama_io_write_host (tensor→backend map, queue reads, sync once). Left out the range-API and server-side streaming parts, since this PR's worker transfers state through the existing llama_state_seq_get_data_ext path - the async reads plug in underneath it transparently.
  • 32MB SO_RCVBUF/SNDBUF socket buffers from my branch (disagg: RPC batching, KV prefetch streaming, sync fence, disagg fixes #5), added next to the existing TCP_NODELAY/SO_REUSEADDR setup in transport.cpp, on both the connect and accept paths.
    All three pieces applied cleanly on top of this PR's branch (shared recent master base), ~840-line patch total, both sides rebuilt together.

Setup: Mac Studio M2 Ultra (Metal, decode) + DGX Spark GB10 (CUDA, prefill) over direct 10GbE, MTU 8000.
Qwen3.6-27B Q4_K_XL, 26.5K-token prompt, 300 tokens out, temp 0. Turn 2 = same conversation + ~250-token continuation, cache_prompt: true.

Config Cold e2e (26.5K prompt) Prompt processing KV transfer (1031 MiB) Decode @26.5K Turn-2 prompt time
ggml-org#25675 stock (blocking, prefill worker) 64.2–65.1 s 42.1–42.9 s inside prompt time 13.5–13.6 t/s 42.6 s — full re-prefill via worker
Hybrid (ggml-org#25675 + async reads + 32MB buffers) 64.8–64.9 s 42.6–42.7 s inside prompt time 13.6 t/s 42.6 s
Hybrid + --prefill-min-tokens 8192 — — — 13.3 t/s 102 s — full local re-prefill on Metal
Streaming branch (#5) 68.1 s 45.4 s 0.5 s 13.2 t/s 2.9 s — extends from cached prefix
decode2 (async range-API streaming) 68.5 s 45.3 s 1.3 s 13.3 t/s —

Decode speed was 13.2–13.6 t/s at that depth across all configs - parity, as expected for memory-bound decode.
Conclusions:

  1. The transport additions changed nothing on cold requests - hybrid is within noise of stock. On a 10GbE link the transfer is simply not the bottleneck of this PR's design; its ~3s advantage over the streaming branches comes from how the prefill itself is organized (one monolithic pass in a dedicated context vs. chunked passes through the shared scheduler), not from how bytes move. So on fast links, blocking transfer is genuinely fine - the earlier concern about it doesn't show up in end-to-end numbers here.
  2. The multi-turn concern is real, though, and it's worse than "re-prefill through the worker". Turn 2 re-prefills the entire 26.7K context via the worker (42.6s). I tried --prefill-min-tokens 8192 expecting the ~250-token continuation to extend locally from cache - instead it recomputed the whole context locally on Metal (102s, which is slower than the worker since GB10 prefills ~2.4x faster than an M2 Ultra). It looks like the prompt cache doesn't survive the worker>target state transfer, so once a conversation has gone through the worker once, there's nothing to extend from. For comparison, a cached-prefix-aware path does turn 2 in 2.9s - a 15–35x difference on exactly the workload (agents, long multi-turn sessions) where disagg matters most.
  3. So the best-of-all-worlds isn't a transport hybrid - it's this PR's worker design for cold prompts + a cache-aware skip for continuations (only route to the worker when the uncached suffix is large, and make sure the transferred state lands in the prompt cache so later turns can extend it).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants