Repository navigation
Conversation
|
This looks like an interesting and simple fix, @ggerganov @ngxson what do you guys think? |
|
I haven't observed unnecessary checkpoint invalidation with recurrent models, so I am not sure what the change is trying to fix. Most of the reports that we get are due to the client injecting stuff in earlier messages which is inefficient, but I don't think we need to try to support. As soon as I observe a valid problem, or get a proper report with a reproduction, I will fix it. I'm using pi daily and haven't observed any problems recently. The only optimization that is currently missing is a follow-up to #22929 to consider past user messages (not just the last one). |
I'm using ryzen 395 max+ and Qwen 3.6 27b, as well as Qwen 3.5 122b. My cache is constantly being flushed, and the processing promt is being created again every time. This fix has resolved the issue. I've tested it on LM Studio on Vulcan (amd radeon 8060s). |
|
Log: |
be5c783 to
750c8d8
Compare
|
@Regrad thank you for this branch. it was driving me crazy that Qwen 3.6 35b but also gemma 4 26b as MoE models were always giving me the message "forcing full prompt re-processing due to lack of cache data (likely due to SWA or hybrid/recurrent memory" in the last few weeks which led me to searching and finding this pr. One thing that I still noticed when running this benchmark for llm's https://github.com/alexziskind1/codeneedle was GLM 5.2 worked on checking this large gap of over 2000 tokens between what was invalidated and came up with this investigation and small change I don't know if it's of any use for your work, if this exposes another "problem" with checkpoints and ub values (mine was -b 8192 and -ub 2048) but it my case I got better cache reuse. The fix or changes might be somewhat wrong, just leaving this out here in case it might make any sense for your work in this pr. |
I have reviewed and added your improvement. Since it was created using LLM, I have checked it and made changes to ensure that your correction only affects hybrid models. |
fc18337 to
67798d1
Compare
9b11fe7 to
6647155
Compare
Ported from upstream ggml-org#24035 items 1-6. These address losing prompt-cache reuse when a client rewrites the prompt ahead of the surviving checkpoints: the checkpoints kept were only the ones past the divergence point, so every turn re-evaluated the whole prompt. - Invalidate a checkpoint only when a token it captured has left the common prefix, instead of whenever pos_max passed the post-restore resume point. Restoring an older checkpoint no longer discards newer valid ones. Applied to both erase sites, including the planned_p0 branch upstream does not have. - Evict from the densest interior of the checkpoint list rather than always dropping the oldest prefix, so the earliest and latest anchors survive the per-slot cap. - Collapse pos_min to pos_max for exact-state memories (FULL / RS seq_rm), resolving the [TAG_CHECKPOINTS_FIX_POS_MIN] TODO for recurrent and hybrid contexts while leaving SWA's under-reported range alone. - Re-capturing an existing boundary re-owns that checkpoint instead of serializing a duplicate, refreshing id_task so the min-step prune exempts it. - Select the runtime restore candidate on the lexical common prefix via server_prompt::find_reusable_checkpoint(), which carries the KVarN descriptor alignment requirement. The covered-range scan remains for removable KV/SWA. - Capture a durable checkpoint at the real common-prefix boundary after an exact-state restore, so the boundary is not lost to the next divergence. The position is snapped to the descriptor alignment. - Add --checkpoint-max-step (-cmx, LLAMA_ARG_CHECKPOINT_MAX_SPACING_NT) to force a checkpoint at least every N tokens so coverage no longer depends on where chat message boundaries happen to fall. Defaults to 0 (off). Spacing lands in [max_step, max_step + n_batch) because checkpoints align to batches. The min-step prune is relaxed to min(min_step, max_step) so the cadence is not undone, and startup warns when --ctx-checkpoints cannot cover the context. The new selection, eviction and cadence helpers are pure functions covered by test-server-prompt-checkpoint. Verified on the CPU build: full ctest suite 124/124. Co-authored-by: CommandCodeBot <noreply@commandcode.ai>
…ersation, so a task's first call takes up a checkpoint at the prefix's end instead of processing 2,220 tokens again (19 s on six CPU cores -> 0.2 s, measured on the lab); S read: the server's checkpoint flags did nothing (S1), drafting is 1.12x on a CPU at n-max 3 (S3) The server makes a context checkpoint only at the end of each prompt it processes, so after a task's first call none lies at the shared prefix and the next task's first call reprocesses it all. Warming the prefix alone leaves the checkpoint where the next conversation needs it: first call 25-72 tokens with cache_n 2220. Launcher-side, because the maintainers keep a new conversation's prefix out of the server's scope. Sources: ggml-org/llama.cpp#22384 ; ggml-org/llama.cpp#24035 ; the measurement is in locallm/PREDICT-2026-10-05-assistant.md (S)
…ersation, so a task's first call takes up a checkpoint at the prefix's end instead of processing 2,220 tokens again (19 s on six CPU cores -> 0.2 s, measured on the lab); S read: the server's checkpoint flags did nothing (S1), drafting is 1.12x on a CPU at n-max 3 (S3) The server makes a context checkpoint only at the end of each prompt it processes, so after a task's first call none lies at the shared prefix and the next task's first call reprocesses it all. Warming the prefix alone leaves the checkpoint where the next conversation needs it: first call 25-72 tokens with cache_n 2220. Launcher-side, because the maintainers keep a new conversation's prefix out of the server's scope. Sources: ggml-org/llama.cpp#22384 ; ggml-org/llama.cpp#24035 ; the measurement is in locallm/PREDICT-2026-10-05-assistant.md (S)
Summary
Preserve reusable prompt checkpoint coverage for recurrent and hybrid models when a follow-up request diverges from previously processed history.
These models require an exact saved memory state. Restoring an older checkpoint previously caused later checkpoints to be invalidated against the restored position, even when their complete token prefix still matched the new request. A short unrelated request could also prevent a much more useful cached conversation from being restored.
What changed
FULLandRSsequence-removal modes as exact-position states and normalize newly captured checkpoints to their actual saved position.Tests
Added
test-server-prompt-cachecoverage for:git diff --checkpasses. The new C++ test has not been compiled or executed as part of this update.Expected effect
Follow-up requests should reprocess only from the newest valid checkpoint before the first changed token. Interleaved short requests should no longer block restoration of a cached long conversation when that conversation provides substantially more reusable checkpoint coverage.
A suffix after the first changed token is not reused because causal/recurrent state depends on the complete preceding prefix.