Skip to content

server: preserve reusable checkpoints for recurrent / hybrid prompts - #24035

Draft
Regrad wants to merge 3 commits into
ggml-org:masterfrom
Regrad:fix/qwen-hybrid-checkpoint-reuse
Draft

Regrad wants to merge 3 commits into
ggml-org:masterfrom
Regrad:fix/qwen-hybrid-checkpoint-reuse

Conversation

@Regrad

@Regrad Regrad commented Jun 2, 2026 •

Copy link
Copy Markdown

Summary

Preserve reusable prompt checkpoint coverage for recurrent and hybrid models when a follow-up request diverges from previously processed history.

These models require an exact saved memory state. Restoring an older checkpoint previously caused later checkpoints to be invalidated against the restored position, even when their complete token prefix still matched the new request. A short unrelated request could also prevent a much more useful cached conversation from being restored.

What changed

  • Classify FULL and RS sequence-removal modes as exact-position states and normalize newly captured checkpoints to their actual saved position.
  • Select exact-state checkpoints by the number of saved tokens contained in the longest common prefix, leaving at least one token for logits.
  • Invalidate checkpoints only when their saved token prefix extends beyond the LCP. Restoring an older checkpoint no longer removes later checkpoints that are still valid.
  • Split prompt batching at a recovered LCP boundary and capture the decoded state there before processing new tokens.
  • Create a checkpoint before entering a media chunk so replacing or adding media does not force processing from the start of the prompt.
  • Retain broad checkpoint coverage by evicting the densest interior checkpoint while preserving early and recent anchors.
  • Rank recurrent/hybrid RAM prompt-cache candidates by the number of tokens that can actually be restored. Cache swapping requires a meaningful improvement derived from half the configured checkpoint spacing.
  • Keep the existing prompt-cache selection behavior for states that support direct range rollback.

Tests

Added test-server-prompt-cache coverage for:

  • selecting the newest checkpoint inside an edited prompt's LCP;
  • direct reuse when tokens are appended to an exact full state;
  • evaluating one token again for an identical prompt;
  • falling back to zero without a valid exact-state checkpoint;
  • retaining direct LCP reuse for non-exact states.

git diff --check passes. The new C++ test has not been compiled or executed as part of this update.

Expected effect

Follow-up requests should reprocess only from the newest valid checkpoint before the first changed token. Interleaved short requests should no longer block restoration of a cached long conversation when that conversation provides substantially more reusable checkpoint coverage.

A suffix after the first changed token is not reused because causal/recurrent state depends on the complete preceding prefix.

@Regrad
Regrad requested review from a team as code owners June 2, 2026 16:52
@Regrad Regrad closed this Jun 2, 2026
@Regrad Regrad reopened this Jun 2, 2026
@pwilkin

pwilkin commented Jun 3, 2026

Copy link
Copy Markdown
Member

This looks like an interesting and simple fix, @ggerganov @ngxson what do you guys think?

@ggerganov

Copy link
Copy Markdown
Member

I haven't observed unnecessary checkpoint invalidation with recurrent models, so I am not sure what the change is trying to fix. Most of the reports that we get are due to the client injecting stuff in earlier messages which is inefficient, but I don't think we need to try to support.

As soon as I observe a valid problem, or get a proper report with a reproduction, I will fix it. I'm using pi daily and haven't observed any problems recently. The only optimization that is currently missing is a follow-up to #22929 to consider past user messages (not just the last one).

@Regrad

Regrad commented Jun 3, 2026

Copy link
Copy Markdown
Author

I haven't observed unnecessary checkpoint invalidation with recurrent models, so I am not sure what the change is trying to fix. Most of the reports that we get are due to the client injecting stuff in earlier messages which is inefficient, but I don't think we need to try to support.

As soon as I observe a valid problem, or get a proper report with a reproduction, I will fix it. I'm using pi daily and haven't observed any problems recently. The only optimization that is currently missing is a follow-up to #22929 to consider past user messages (not just the last one).

I'm using ryzen 395 max+ and Qwen 3.6 27b, as well as Qwen 3.5 122b. My cache is constantly being flushed, and the processing promt is being created again every time. This fix has resolved the issue. I've tested it on LM Studio on Vulcan (amd radeon 8060s).

@Regrad

Regrad commented Jun 3, 2026

Copy link
Copy Markdown
Author

Log:

18: slot update_slots: id  3 | task 17322 | new prompt, n_ctx_slot = 262144, n_keep = 15, task.n_tokens = 15
19: slot update_slots: id  3 | task 17322 | cache reuse is not supported - ignoring n_cache_reuse = 256
20: slot update_slots: id  3 | task 17322 | n_past = 15, slot.prompt.tokens.size() = 22, seq_id = 3, pos_min = 21, n_swa = 1
21: slot update_slots: id  3 | task 17322 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
22: slot update_slots: id  3 | task 17322 | n_tokens = 0, memory_seq_rm [0, end)
23: slot update_slots: id  3 | task 17322 | prompt processing progress, n_tokens = 11, batch.n_tokens = 12, progress = 0.733333
24: [2026-04-02 11:55:04][INFO][qwen3.5-122b-a10b@?] Prompt processing progress: 0.0%
192407: slot update_slots: id  0 | task 97964 | forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory, see https://github.com/ggml-org/llama.cpp/pull/13194#issuecomment-2868343055)
192408: slot update_slots: id  0 | task 97964 | erased invalidated context checkpoint (pos_min = 154657, pos_max = 154657, n_tokens = 154658, n_swa = 0, pos_next = 0, size = 62.813 MiB)
192409: [2026-05-02 23:14:58][DEBUG] slot update_slots: id  0 | task 97964 | erased invalidated context checkpoint (pos_min = 154767, pos_max = 154767, n_tokens = 154768, n_swa = 0, pos_next = 0, size = 62.813 MiB)
192410: [2026-05-02 23:14:58][DEBUG] slot update_slots: id  0 | task 97964 | erased invalidated context checkpoint (pos_min = 155123, pos_max = 155123, n_tokens = 155124, n_swa = 0, pos_next = 0, size = 62.813 MiB)

194477: [2026-05-02 23:23:01][DEBUG] slot update_slots: id  0 | task 98398 | restored context checkpoint (pos_min = 18818, pos_max = 18818, n_tokens = 18819, n_past = 18819, size = 62.813 MiB)
195051: [2026-05-02 23:23:30][DEBUG] slot update_slots: id  0 | task 98572 | restored context checkpoint (pos_min = 8191, pos_max = 8191, n_tokens = 8192, n_past = 8192, size = 62.813 MiB)

@Regrad
Regrad marked this pull request as draft June 4, 2026 09:19
@Regrad
Regrad force-pushed the fix/qwen-hybrid-checkpoint-reuse branch from be5c783 to 750c8d8 Compare June 15, 2026 11:25
@Regrad
Regrad marked this pull request as ready for review June 15, 2026 14:30
@ichim-david

ichim-david commented Jun 15, 2026 •

Copy link
Copy Markdown

@Regrad thank you for this branch. it was driving me crazy that Qwen 3.6 35b but also gemma 4 26b as MoE models were always giving me the message "forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory" in the last few weeks which led me to searching and finding this pr.

One thing that I still noticed when running this benchmark for llm's https://github.com/alexziskind1/codeneedle was
something like this:
3.11.883.089 W slot update_slots: id 0 | task 986 | restored context checkpoint (pos_min = 9675, pos_max = 9675, pos_end = 9676, n_tokens = 9676, n_past = 9676, size = 81.896 MiB) 3.11.883.092 W slot update_slots: id 0 | task 986 | erased invalidated context checkpoint (pos_min = 11718, pos_max = 11718, pos_end = 11719, n_tokens = 11719, n_swa = 0, prefix_end = 11586, pos_next = 9676, size = 85.925 MiB)

GLM 5.2 worked on checking this large gap of over 2000 tokens between what was invalidated and came up with this investigation and small change
https://gist.github.com/ichim-david/b1f635868d62442894caf019041cdaf3

I don't know if it's of any use for your work, if this exposes another "problem" with checkpoints and ub values (mine was -b 8192 and -ub 2048) but it my case I got better cache reuse.

The fix or changes might be somewhat wrong, just leaving this out here in case it might make any sense for your work in this pr.

@Regrad

Regrad commented Jun 15, 2026

Copy link
Copy Markdown
Author

@Regrad thank you for this branch. it was driving me crazy that Qwen 3.6 35b but also gemma 4 26b as MoE models were always giving me the message "forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory" in the last few weeks which led me to searching and finding this pr.

I have reviewed and added your improvement. Since it was created using LLM, I have checked it and made changes to ensure that your correction only affects hybrid models.

@Regrad
Regrad force-pushed the fix/qwen-hybrid-checkpoint-reuse branch from fc18337 to 67798d1 Compare June 15, 2026 20:34
@Regrad
Regrad marked this pull request as draft June 16, 2026 07:20
@Regrad
Regrad marked this pull request as ready for review June 17, 2026 07:49
@Regrad
Regrad requested a review from ggerganov as a code owner June 17, 2026 07:49
@Regrad
Regrad force-pushed the fix/qwen-hybrid-checkpoint-reuse branch from 9b11fe7 to 6647155 Compare July 26, 2026 18:33
@github-actions github-actions Bot added the testing Everything test related label Jul 26, 2026
@Regrad Regrad changed the title server: avoid unnecessary checkpoint invalidation for recurrent / hybrid models server: preserve reusable checkpoints for recurrent / hybrid prompts Jul 26, 2026
@Regrad
Regrad marked this pull request as draft July 27, 2026 09:35
Encoded404 added a commit to Encoded404/beellama-QoL-v100.cpp that referenced this pull request Sep 20, 2026
Ported from upstream ggml-org#24035 items 1-6. These address losing prompt-cache reuse
when a client rewrites the prompt ahead of the surviving checkpoints: the
checkpoints kept were only the ones past the divergence point, so every turn
re-evaluated the whole prompt.

- Invalidate a checkpoint only when a token it captured has left the common
  prefix, instead of whenever pos_max passed the post-restore resume point.
  Restoring an older checkpoint no longer discards newer valid ones. Applied to
  both erase sites, including the planned_p0 branch upstream does not have.
- Evict from the densest interior of the checkpoint list rather than always
  dropping the oldest prefix, so the earliest and latest anchors survive the
  per-slot cap.
- Collapse pos_min to pos_max for exact-state memories (FULL / RS seq_rm),
  resolving the [TAG_CHECKPOINTS_FIX_POS_MIN] TODO for recurrent and hybrid
  contexts while leaving SWA's under-reported range alone.
- Re-capturing an existing boundary re-owns that checkpoint instead of
  serializing a duplicate, refreshing id_task so the min-step prune exempts it.
- Select the runtime restore candidate on the lexical common prefix via
  server_prompt::find_reusable_checkpoint(), which carries the KVarN descriptor
  alignment requirement. The covered-range scan remains for removable KV/SWA.
- Capture a durable checkpoint at the real common-prefix boundary after an
  exact-state restore, so the boundary is not lost to the next divergence. The
  position is snapped to the descriptor alignment.
- Add --checkpoint-max-step (-cmx, LLAMA_ARG_CHECKPOINT_MAX_SPACING_NT) to force
  a checkpoint at least every N tokens so coverage no longer depends on where
  chat message boundaries happen to fall. Defaults to 0 (off). Spacing lands in
  [max_step, max_step + n_batch) because checkpoints align to batches. The
  min-step prune is relaxed to min(min_step, max_step) so the cadence is not
  undone, and startup warns when --ctx-checkpoints cannot cover the context.

The new selection, eviction and cadence helpers are pure functions covered by
test-server-prompt-checkpoint.

Verified on the CPU build: full ctest suite 124/124.

Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
trestoncuzzort added a commit to trestoncuzzort/dawnr that referenced this pull request Oct 6, 2026
…ersation, so a task's first call takes up a checkpoint at the prefix's end instead of processing 2,220 tokens again (19 s on six CPU cores -> 0.2 s, measured on the lab); S read: the server's checkpoint flags did nothing (S1), drafting is 1.12x on a CPU at n-max 3 (S3)

The server makes a context checkpoint only at the end of each prompt it processes, so after a task's first call
none lies at the shared prefix and the next task's first call reprocesses it all. Warming the prefix alone leaves
the checkpoint where the next conversation needs it: first call 25-72 tokens with cache_n 2220. Launcher-side,
because the maintainers keep a new conversation's prefix out of the server's scope.

Sources: ggml-org/llama.cpp#22384 ; ggml-org/llama.cpp#24035 ;
the measurement is in locallm/PREDICT-2026-10-05-assistant.md (S)
trestoncuzzort added a commit to trestoncuzzort/dawnr that referenced this pull request Oct 6, 2026
…ersation, so a task's first call takes up a checkpoint at the prefix's end instead of processing 2,220 tokens again (19 s on six CPU cores -> 0.2 s, measured on the lab); S read: the server's checkpoint flags did nothing (S1), drafting is 1.12x on a CPU at n-max 3 (S3)

The server makes a context checkpoint only at the end of each prompt it processes, so after a task's first call
none lies at the shared prefix and the next task's first call reprocesses it all. Warming the prefix alone leaves
the checkpoint where the next conversation needs it: first call 25-72 tokens with cache_n 2220. Launcher-side,
because the maintainers keep a new conversation's prefix out of the server's scope.

Sources: ggml-org/llama.cpp#22384 ; ggml-org/llama.cpp#24035 ;
the measurement is in locallm/PREDICT-2026-10-05-assistant.md (S)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

examples server testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants