Repository navigation
server: fix checkpoints creation - #22929
Conversation
|
tested following way:
Details |
|
Hi @jacekpoplawski, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
d878621 to
ea9369c
Compare
|
Details |
|
Yes, that seems in a good direction. Have you done testing that it works as expected? |
|
This needs autoparser dedicated support for split-marker detection; currently, this will assume that all autoparser models use the ChatML markers ( I'll try to submit the marker detection code ASAP. |
| @@ -600,6 +600,34 @@ task_params server_task::params_from_json_cmpl( | |||
| throw std::runtime_error("n_cmpl cannot be greater than the number of slots, please increase -np"); | |||
| } | |||
|
|
|||
| const auto message_spans = json_value(data, "message_spans", json::array()); | |||
| if (message_spans.is_array()) { | |||
| int32_t last_user_pos = -1; | |||
There was a problem hiding this comment.
You can probably use 0 as the sentinel value here, since a checkpoint at pos 0 isn't useful. Should help clean up the other logic too.
|
|
||
| if ((size_t) last_user_pos <= prompt.size()) { | ||
| const std::string prefix = prompt.substr(0, (size_t) last_user_pos); | ||
| const auto prefix_tokens = common_tokenize(vocab, prefix, true, true); |
There was a problem hiding this comment.
Just a guess, but this will probably create incorrect checkpoints for multimodal models with at least one image in the prompt.
There was a problem hiding this comment.
Yes, you are right, this breaks after the first image.
There was a problem hiding this comment.
now it should be ok
It works stable for my usecase: pi, qwen 3.6 27B, 200k ctx, 24 checkpoints With 8 checkpoints I was able to reproduce As @aldehir pointed out, this does not work correctly with multimodal prompts. I committed a fallback to the old mechanism for that case. Should I add a switch to enable this new mechanism as an option, or should I try to support multimodal prompts as well? I understand that the impact of this change is significant, but the benefits are also significant: agentic coding is much more responsive now. |
|
Tested, model used: https://huggingface.co/unsloth/Qwen3.6-27B-GGUF The message But there is a cache miss that happens only one time, and I can't reproduce it. I tried for 1h without being able to hit that again. Great work! You saved my time and my electricity bill. Edit: |
| // stop the prompt batch exactly before the latest user input, so a checkpoint | ||
| // can be created at the conversation boundary | ||
| if (checkpoint_before_last_user_token > 0 && | ||
| slot.prompt.n_tokens() == checkpoint_before_last_user_token) { |
There was a problem hiding this comment.
Just for testing, can you check that this works also with images:
| slot.prompt.n_tokens() == checkpoint_before_last_user_token) { | |
| slot.prompt.get_text_tokens().size() == checkpoint_before_last_user_token) { |
And also the same change applied below on line 2752
There was a problem hiding this comment.
I tried using slot.prompt.tokens.get_text_tokens().size(), but it didn’t help. I am now testing an approach that iterates over the full token sequence while skipping LLAMA_TOKEN_NULL, and it seems to work.
|
Just pulled that PR. When using Pi or OpenCode I still get |
Could you say a bit more and show some logs? I tested this fix for many hours (Qwen 3.6 27B), and |
|
@ggerganov Multimodal prompts are now supported. I will continue working on this, because the checkpoint is currently created only after the image is read for the second time, not the first time. If you are happy with the direction of my changes, I’d like to add two arguments: one to set |
I was using qwen3.5 122b. I sadly cant provide more logs as of right now. |
|
@jacekpoplawski Is the |
If the context needs to be recalculated, a checkpoint could be generated for each user message. |
|
我现在使用如下策略来适应新的KVcache逻辑:
I am currently adopting the following strategy to adapt to the new KVcache logic:
|
* common : add common_chat_split_by_role * cont : fix spans to reach end of message * server: fix checkpoints creation - extract message_spans from chat templates - find the prompt token position before the latest user message - split prompt batching at that position - create a context checkpoint before the latest user input - avoid periodic mid-prompt checkpoints when that position is known - handle multimodal prompts when mapping text/template positions to server prompt tokens - add --checkpoint-min-step to control minimum spacing between checkpoints * cont : clean-up * Support autoparser detection for message barriers * server: fix message span delimiter and update docs --------- Co-authored-by: Alde Rojas <hello@alde.dev> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com>
git merge's 3-way resolver did not flag two semantic duplicates in
tools/server/server-context.cpp because the merge base did not contain
either symbol. The duplicate bodies are byte-identical, so removing
the second copy of each pair is semantically equivalent.
Removed:
- `const bool near_prompt_end` declaration at line 3653 (upstream
side, e2ef8fe, "server: fix checkpoints creation" PR ggml-org#22929).
- `static uint32_t server_n_outputs_max(...)` body at lines 219-232
(upstream side, de6f727, "llama: limit max outputs of
llama_context" PR ggml-org#23861; one line modified by 5dcb711,
"speculative: fix n_outputs_max and remove draft-simple auto-enable"
PR ggml-org#23988).
Kept:
- The cache-side copies (72cfbcd), which match the
cache-optimization chain the Stage 11 work is built on.
This comment was marked as low quality.
This comment was marked as low quality.
|
This still is an issue, using opencode I constantly get "forcing full prompt re-processing due to lack of cache data" when it is doing agentic stuff. The only time it doesn't seem to do it is with simple follow up questions. These are my settings. If I can do anything to provide more information feel free to ask. |
please run llama-server with |
Such a message seems to be reassuring and stable in terms of reusing the kv cache: |
|
最新发生了一个变更,引用了22929的common_chat_split_by_role,需要增加一个相关pr的回滚。 git revert aedb2a5e9 --no-editA recent change has been made, which references common_chat_split_by_role from 22929. A rollback for the related PR is required. git revert aedb2a5e9 --no-edit |
|
tag:b9655 是最后一个简单运行revert可以成功编译的。 tag:b9656 (#24329) 修改的内容有冲突,revert已经不能自动处理冲突了。抽空得研究一下如何处理~~~ Tag b9655 is the last one where a simple git revert results in a successful compilation. Tag b9656 (#24329) introduces conflicting changes, so git revert can no longer automatically resolve the conflicts. I need to look into how to handle this when I have some time. |
|
鉴于已经不好revert了。 Given that reverting is no longer feasible, |
|
最新情况报告,前端APP120秒超时重新提交对话,能够正常继续处理。 |
|
Hitting the same issue with openclaw + qwen3.6 . If I understood correctly, the issue happens because this PR changes checkpoints creation to be message boundary-based, and since some API consumers (openclaw) change the system prompt that is supposed to be static, the whole cache is trashed. |
|
Maybe there needs to be a new issue reported as this is "Merged"? |
This reverts commit e2ef8fe. Assisted-by: synthmerge
* common : add common_chat_split_by_role * cont : fix spans to reach end of message * server: fix checkpoints creation - extract message_spans from chat templates - find the prompt token position before the latest user message - split prompt batching at that position - create a context checkpoint before the latest user input - avoid periodic mid-prompt checkpoints when that position is known - handle multimodal prompts when mapping text/template positions to server prompt tokens - add --checkpoint-min-step to control minimum spacing between checkpoints * cont : clean-up * Support autoparser detection for message barriers * server: fix message span delimiter and update docs --------- Co-authored-by: Alde Rojas <hello@alde.dev> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com>
* common : add common_chat_split_by_role * cont : fix spans to reach end of message * server: fix checkpoints creation - extract message_spans from chat templates - find the prompt token position before the latest user message - split prompt batching at that position - create a context checkpoint before the latest user input - avoid periodic mid-prompt checkpoints when that position is known - handle multimodal prompts when mapping text/template positions to server prompt tokens - add --checkpoint-min-step to control minimum spacing between checkpoints * cont : clean-up * Support autoparser detection for message barriers * server: fix message span delimiter and update docs --------- Co-authored-by: Alde Rojas <hello@alde.dev> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com>
This reverts commit e2ef8fe. Assisted-by: synthmerge Synthmerge status: Resolved: tools/server/server-context.cpp Synthmerge status: Resolved: common/common.h tools/server/server-context.cpp
This reverts commit e2ef8fe. Assisted-by: synthmerge Synthmerge status: Resolved: tools/server/server-context.cpp Synthmerge status: Resolved: common/common.h tools/server/server-context.cpp
* common : add common_chat_split_by_role * cont : fix spans to reach end of message * server: fix checkpoints creation - extract message_spans from chat templates - find the prompt token position before the latest user message - split prompt batching at that position - create a context checkpoint before the latest user input - avoid periodic mid-prompt checkpoints when that position is known - handle multimodal prompts when mapping text/template positions to server prompt tokens - add --checkpoint-min-step to control minimum spacing between checkpoints * cont : clean-up * Support autoparser detection for message barriers * server: fix message span delimiter and update docs --------- Co-authored-by: Alde Rojas <hello@alde.dev> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: Piotr Wilkin <piotr.wilkin@syndatis.com>

Overview
Implemented as requested in #22826 (comment)
message_spansfrom chat templatesAdditional information
This is another chapter in my journey toward fixing
forcing full prompt re-processing due to lack of cache dataMy main goal is to increase the "responsiveness" of agentic coding in llama.cpp
I am currently testing this with the following command:
preserve_thinkingreally helps, without it, the prompt history changes, so there is always some reprocessingRequirements