Repository navigation
Conversation
Spread the disaggregated prefill over update_slots ticks instead of blocking the loop. Add SLOT_STATE_DISAGG_PREFILL and a per slot cursor, split disagg_prefill into begin/step/finish, and step one chunk per tick so other slots decode on ctx_tgt between chunks. The handoff and the last token run on return via the existing prompt tail. Reset n_empty_consecutive on prefill progress so the empty batch watchdog does not abort while a lone slot prefills on ctx_pfx. Output is token for token identical to local decode.
Replace the single fixed prefill sequence with a pool of n_prefill sequences on ctx_pfx. A slot acquires a free sequence when it enters disaggregated prefill and releases it on handoff, so up to n_prefill prefills run concurrently, each on its own sequence. When the pool is full the slot falls back to local prompt processing. Move the per chunk trace to SLT_DBG and drop its timing. The per request summary stays at SLT_INF.
The disaggregated prefill path does not hand off the draft context KV, so it is skipped when a draft context is active. Speculative decoding and disagg prefill are mutually exclusive for now. Document --prefill-device, --n-prefill and --n-prefill-chunk in the server README, with a disaggregated prefill/decode section and an RPC example.
The disagg prefill path went through the normal prompt tail with n_past = 0, which dropped the prefix and erased the KV just handed off into ctx_tgt, so the prompt was reprocessed locally and the remote prefill was wasted. Set n_past to the prefilled prefix length and skip the checkpoint/pos_min reset for disagg, so ctx_tgt only evaluates the last token. Split the slot flag into disagg_attempted (try once, no retry loop) and disagg_prefilled (handoff valid), set the latter only after the state get/set succeed, check their returned sizes and fall back to local processing on failure. Clamp n-prefill-chunk to at least 1.
disagg_prefill_batch packs one chunk from every prefilling slot into a single llama_decode on ctx_pfx, replacing the per slot disagg_prefill_step that issued one decode per slot. it runs once before the prompt batching loop, and slots that complete their prefix flip to SLOT_STATE_STARTED so the KV handoff and the final token run on ctx_tgt the same tick. a failed ctx_pfx decode falls every batched slot back to local prompt processing: disagg_pos stays put so disagg_prefill_finish releases the sequence without a handoff and disagg_prefilled stays false.
gate disagg prefill on SERVER_TASK_TYPE_COMPLETION so embedding and rerank keep their full ctx_tgt outputs, and skip it when the slot carries lora adapters since ctx_pfx never sets them and a lora prefix KV would be wrong. in disagg_prefill_batch only the slots packed into the ctx_pfx batch advance and fall back on a failed decode, guarded against an empty batch.
create ctx_pfx with the same n_seq_max as ctx_tgt so both contexts share the same n_stream, otherwise llama_state_seq_set_data rejects the handoff with an n_stream mismatch and every prefill silently falls back to local. clamp n_prefill to n_parallel since the pool cannot hand off more concurrent prefills than there are decode slots.
Six independent patches on top of the disaggregated prefill/decode server, developed and benchmarked running Qwen3.6-27B/35B-A3B across a direct 10GbE link (Mac Studio M2 Ultra decode <-> DGX Spark GB10 CUDA prefill): - process_mtmd_prompt: has_mtmd = !files.empty(), so a loaded --mmproj no longer unconditionally disables disagg prefill for text-only requests. - server-context: skip the disagg path when the slot already has a cached common prefix, so multi-turn conversations extend the local KV instead of re-shipping the whole prefix through the remote prefill device on every turn. - New RPC_CMD_GET_TENSOR_BATCH (proto v4.1): batches the post-prefill KV fetch into one round trip instead of one per tensor. Raised measured wire throughput on our 10GbE link from ~3.8 to ~9.8 Gbps. - transport: 32MB SO_RCVBUF/SNDBUF on RPC sockets, independent win on top of the batching above. - ggml_backend_rpc_synchronize was a no-op (graph_compute over RPC is fire-and-forget). Made it issue a cheap round-trip command as a real fence -- flagging this as a correctness fix, not just perf: a caller chaining multiple RPC graphs (e.g. disagg prefill chunks) could previously read a buffer before queued remote work had completed. - KV prefetch streaming during prefill: llama_state_seq_prefetch_ext -> llama_memory_i::state_prefetch (implemented for llama_kv_cache, forwarded through llama_memory_hybrid -- required for Qwen3.6's hybrid recurrent/attention architecture) -> a fire-and-forget ggml_backend_rpc_prefetch_tensor_batch request, drained lazily into a client-side region cache keyed by (remote tensor ptr, offset). Hides the attention-KV transfer for each prefill chunk behind the next chunk's remote compute. Measured: 28.8K-token prompt KV handoff 1264ms -> 544ms (86% served from prefetch cache). See PR description for full benchmark tables (cluster vs. single-node, depth scaling, speculative decoding interaction, layer-split mode) and deployment notes.
…I config dialog with slider Assisted-by: Claude
…ements # Conflicts: # tools/server/server-context.cpp # tools/server/server-models.cpp # tools/server/server-models.h # tools/ui/src/lib/components/app/models/ModelsSelectorOption.svelte # tools/ui/src/lib/constants/api-endpoints.ts # tools/ui/src/lib/services/models.service.ts
|
Interesting; I'm going to rebase my branch to see more ! |
f4a0744 to
429802c
Compare
|
I've rebased on my end. I'm going to rebase yours to see. |
|
OK the prefetch path looks like a hack to me, we need real async primitives in libllama for this. I rebased your commit here: https://github.com/ServeurpersoCom/llama.cpp/tree/disagg-rpc-improvements / ggml-org@35dc99d |
|
Yes, all numbers are from real hardware: Mac Studio M2 Ultra (192GB) + DGX Spark GB10, direct 10GbE link, MTU 8000. Happy to re-run any specific benchmark if useful |
|
The official discussion can be found here : ggml-org#21266 |
|
ran your decode2 on the same hardware - e2e parity with my branch (45.5s prompt / 13.3 tok/s decode on a 26.5K prompt), but your get(net) is 1.3s vs 0.5s on mine - the 32MB socket buffers alone should close most of that gap |
|
Look this one : ggml-org#25675 We shouldn't hesitate to test all possible implementations in order to extract the best approach for llama.cpp even if it means combining different ideas! Also, you really shouldn’t trust agentic benchmarks without asking for a simple, reproducible test script and running it yourself. Things can quickly go off the rails: the cobra effect is always hidden and full of surprises. |
|
tested all three on the same rig (M2 Ultra <> GB10, 10GbE, 26.5K prompt): #25675 blocking = 64-65s e2e, my streaming = 68.1s, decode2 = 68.5s. On a fast link blocking actually wins - the dedicated prefill worker beats chunked scheduling by more than streaming saves. On 1Gbps the math flips: 1GB KV = ~10s serialized per request. Suggests: #25675's worker design + streaming as an option for slow links + the multi-turn cache skip |
|
It's still better to work on upstream (ggml-org) than on my fork! |
|
Well, it was a good try What the hybrid was: Base: this PR (ggml-org#25675) - the dedicated prefill worker design
Setup: Mac Studio M2 Ultra (Metal, decode) + DGX Spark GB10 (CUDA, prefill) over direct 10GbE, MTU 8000.
Decode speed was 13.2–13.6 t/s at that depth across all configs - parity, as expected for memory-bound decode.
|
Hi, thanks for the disagg prefill/decode work — I've been running it on a direct 10GbE link between a Mac Studio (M2 Ultra, Metal decode) and a DGX Spark (GB10, CUDA prefill), on Qwen3.6-27B and Qwen3.6-35B-A3B. This PR is a set of patches on top of
poc/server-disagg-prefill-decode(based on f4a0744), plus the numbers I measured along the way, in case any of it is useful upstream.Patches
has_mtmd = !files.empty()inprocess_mtmd_prompt(server-common.cpp) — as written, having an--mmprojloaded unconditionally disabled disagg prefill even for pure-text requests. This mattered a lot since vision models are common now (Qwen3.6-VL etc.) and I didn't want to give up disagg just because the model can see images.Skip disagg when the slot has a cached common prefix (server-context.cpp) — disagg re-establishes the whole prefix KV on every request by design, which is right for a cold prompt but wasteful for multi-turn conversations where most of the prefix is already resident locally. Added a check on
get_common_prefix()before triggering the disagg path, so follow-up turns extend the local KV instead of re-shipping everything through the remote prefill device.Batched RPC tensor fetch — new
RPC_CMD_GET_TENSOR_BATCH(proto v4.1). The KV handoff after disagg prefill was doing one round-trip per tensor, which capped wire throughput at ~3.8 Gbps on our 10GbE link regardless of MTU/buffer tuning. Batching the fetch into one round trip got us to ~9.8 Gbps (essentially wire speed).32MB SO_RCVBUF/SNDBUF on the RPC sockets (transport.cpp) — default OS buffers were bottlenecking the bulk tensor transfers independent of (3).
ggml_backend_rpc_synchronizeis currently a no-op —graph_computeover RPC is fire-and-forget (no response expected), so there's nothing to block on. I think this is worth flagging as a correctness issue, not just a perf one: on a slot doing multiple RPC graphs back to back (e.g. disagg prefill chunks), a caller relying onsynchronizeas a barrier can read a buffer before the queued work has actually completed on the remote side. I made it issue a cheap round-trip command (DEVICE_COUNT) as a real fence — since the remote processes commands in order, getting a response means the queue up to that point has drained.KV prefetch streaming during prefill — this was the biggest win. Added
llama_state_seq_prefetch_ext→llama_memory_i::state_prefetch(implemented forllama_kv_cache, forwarded throughllama_memory_hybrid— Qwen3.6 is a hybrid recurrent/attention architecture, so the forwarding turned out to be required, not optional) →ggml_backend_rpc_prefetch_tensor_batch, a fire-and-forget prefetch request whose response is drained lazily into a client-side region cache keyed by (remote tensor ptr, offset). Called per-chunk insidedisagg_prefill_batch, with the per-chunkllama_synchronizeremoved since queue ordering is enough to guarantee correctness. Net effect: the attention-KV transfer for a chunk is fully hidden behind the next chunk's remote compute, instead of happening as a separate blocking step after prefill finishes.Numbers (Qwen3.6-27B UD-Q4_K_XL, Mac Studio M2 Ultra ↔ DGX Spark GB10, direct 10GbE, MTU 8000)
KV handoff, 28.8K-token prompt: 1264ms → 544ms after patch 6 (86% of the ~1.1GB transfer served from the prefetch cache; the residual ~150ms is the fixed 149.6MiB recurrent DeltaNet state, which is not prefetchable, plus the final chunk's compute).
End-to-end, same 28.8K-token prompt: prefill 48.9s on Spark (~590 tok/s) + 0.57s handoff + 31.3s decode on Mac (~12.8 tok/s at that depth).
Cluster vs. single-node (same ~2200-token prompt, temp=0, same model):
--device RPC0)Spark-only decode being ~2.7x slower than Mac isn't a disagg issue — it's GB10 being memory-bandwidth-bound on single-batch autoregressive decode, consistent with what others have reported for this hardware. The disaggregated split plays to each device's strength: CUDA for parallel prefill, Metal for decode.
Depth scaling (
llama-bench, tg128, cluster prefill device):Speculative decoding (
--spec-type ngram-mod) stacks fine with disagg, but with a caveat: on the first disagg request in a session, draft-accept rate is low (observed ~9%) because the CUDA prefill and Metal decode produce numerically slightly different logits at temp=0 (verified — outputs diverge starting a few dozen tokens in on identical greedy-decode requests), so n-grams learned from one don't transfer cleanly to the other on that first pass. On follow-up turns disagg is skipped anyway (see patch 2), so ngram-mod runs at full effectiveness there. Net gain on repetitive/code-edit workloads: +39% (24→33.4 tok/s) on a code-edit task, up to 2.1x on highly repetitive prompts, neutral (no regression) on fresh text.Layer-split mode (
--device MTL0,RPC0 --split-mode layer --tensor-split, no--prefill-device) for models too large for either machine alone — tested on Qwen3.6-35B-A3B-Q8 (34GB) split 55/45: decode 18.6 tok/s vs. 65.6 tok/s running the same quant Mac-only. Expected — every token's activations now cross the 10GbE link — but it's the only way to run something that genuinely doesn't fit in 147GB (Mac) or 119GB (Spark) alone, giving ~260GB combined. Not something I'd use unless the model actually requires it.Network / deployment notes (not code changes, but cost me real debugging time)
net.core.wmem_max/rmem_max= 16MB andnet.ipv4.tcp_slow_start_after_idle=0on the Spark side (persisted via/etc/sysctl.d/),kern.ipc.maxsockbuf=16777216on the Mac (/etc/sysctl.conf) plusnetworksetup -setMTU. Withouttcp_slow_start_after_idle=0specifically, throughput on a request after any idle gap re-ramps from a slow-start window instead of going straight to line rate — noticeable since our traffic pattern is bursty (idle between requests, then a burst per prefill).llama-serverbinary run as a normal user) that tries toconnect()to the RPC port getserrno 65("No route to host") even though the host is reachable —nc/system tools connect fine, confirming it's an app-level Local Network permission gate, not routing. There's no user-facing prompt for this in our testing; it just fails. The workaround we're running with is launchingllama-serverfrom a root-owned LaunchDaemon (root is exempt from the check), but that's a workaround, not a fix — a note in the README about this would save the next Mac user a confusing debugging session, since the error message points at routing/firewall, not permissions.majorequal / serverminor<= client, so a stalerpc-serveron the remote box fails closed rather than degrading, which is good, but easy to trip over when only rebuilding one side.rpc-serveron the CUDA box under asystemdunit (Restart=always) and the routerllama-serveron macOS under the LaunchDaemon mentioned above — both survive a reboot and a crash without manual intervention. Worth doing before treating this as more than a demo, since a single stuck RPC connection previously required killing and restarting both processes.Scope of this PR
Self-contained on top of f4a0744, ~15 files touched across
ggml/src/ggml-rpc/,src/llama-kv-cache.{cpp,h},src/llama-memory-hybrid.{cpp,h}, andtools/server/server-context.cpp. Happy to split it up by concern if you'd rather review RPC batching / prefetch streaming / disagg fixes separately — they're fairly independent of each other.