Skip to content

chat : add message delimiters to the DeepSeek V3.2/V4 parser - #29008

Merged
pwilkin merged 2 commits into
ggml-org:masterfrom
midagedev:deepseek-msg-delimiters
Sep 17, 2026
Merged

pwilkin merged 2 commits into
ggml-org:masterfrom
midagedev:deepseek-msg-delimiters

Conversation

@midagedev

Copy link
Copy Markdown
Contributor

Assisted-by: Claude

Overview

안녕하세요. 영어를 잘 하지 못해서 한국어로 쓰고, 아래에 번역을 붙입니다.

DeepSeek V4 Flash를 코딩 에이전트에서 쓰다가 프리필이 지나치게 오래 걸리는 문제를 만났습니다. 분석해 보니 다른 모델과 달리 DeepSeek 전용 파서가 유저 메시지 구분자를 서버에 알려주지 않아서, 유저 턴 위치에 체크포인트가 생기지 않고 프롬프트 캐시가 거의 활용되지 않는 것이 원인이었습니다.

포크해서 파서에 구분자 두 개를 추가했더니 캐시 적중이 명확하게 올랐습니다. #24176이 같은 처리를 다른 파서들에 넣었는데, 그 시점에 이미 있던 DeepSeek V3.2 파서는 목록에 빠져 있었습니다.

검증 자료를 아래에 첨부했으니 검토 부탁드립니다.


(English, translated from the Korean above)

Hello. My English is not good, so I wrote this in Korean; a translation follows.

While using DeepSeek V4 Flash with a coding agent, prefill took far too long. The cause turned out to be different from other models: the DeepSeek specialized parser does not report the user message delimiters to the server, so no checkpoint is created at user turns and the prompt cache is barely used.

I added the two delimiters to the parser in a fork and cache reuse went up clearly. #24176 added the same thing to the other parsers, but the DeepSeek V3.2 parser, which already existed at that time, was not in that list.

Verification data is attached below.

Change: common_chat_params_init_deepseek_v3_2 now sets data.message_delimiters to USER <|User|> and ASSISTANT <|Assistant|>. Both markers open every turn in the V3.2, V4 and V4.1 templates.

Test: test_deepseek_message_delimiters in tests/test-chat.cpp, V3.2 and V4 templates. Fails on master (Expected: <|User|> Actual: empty, rc 134), passes with the patch. Same style as test_msg_token_delimiters_split next to it; no new test file.

Additional information

Why it matters for this model: DeepSeek V3.2/V4/V4.1 use SWA layers (n_swa 128, llama_kv_cache_dsv4). The KV cache cannot be rolled back to a position without a checkpoint, and without delimiters the server only places checkpoints at end - 4 and end - 4 - n_ubatch. So any request that diverges before the tail re-prefills from token 0.

Setup: DeepSeek-V4.1-Flash, one RTX A6000, experts on host (-ot "exps=CPU"), llama-server default checkpoint settings. Prompt = first request of a coding agent, 13 167 tokens (4 200 system, ~8 950 tool schemas for 23 tools, 20 user). Prefill ~46 tok/s. Numbers from timings in the response.

master (checkpoints in the log: n_tokens 12 651 and 13 163 only):

change in the replayed request first differing token cache_n prompt_n wall
none - 13 163 4 0.6 s
new user question 13 146 12 651 510 13.8 s
one word at the end of the system prompt 4 200 0 13 168 236 s
one word in one tool description ~10 400 0 13 170 227 s
same edits with --ctx-checkpoints 64 --checkpoint-min-step 1024 0 full

patched (third checkpoint at n_tokens 13 145 = first user message):

change cache_n prompt_n wall master
new user question 13 145 16 0.8 s 12 651 / 510 / 13.8 s
three user turns 13 145 75 5.5 s -
second of three user turns edited 13 145 79 5.0 s 0 / full

Not tested:

  • V3.2 on a real server (template only, in test-chat)
  • multiple slots
  • speculative decoding on

Unchanged by this PR: an edit inside the system prompt or the tool list still re-prefills from 0. That is the mid-prompt checkpoint policy from #22929, not the parser.

Same omission, not touched here: common/parsers/lfm2.cpp (LFM2 is recurrent + SWA, so it should behave the same; user marker would be <|im_start|>user\n). ministral3, gigachat_v3, kimi_k2, functionary_v3_2 also lack delimiters but serve full-KV models where no checkpoint is needed.

Related: #24176, #21785 (the V3.2 parser), #25452 (DSV4-Flash divergent turns re-prefill from the checkpoint boundary), #21831, #24055, #28302.

Searched before opening: gh pr list --search "message_delimiters", "deepseek delimiters", "deepseek checkpoint"; gh issue list --search "deepseek prompt cache", "V4 checkpoint", "cache_n deepseek"; open PRs touching common/parsers/deepseek.cpp: #28612, #28724 (unrelated).

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. The Korean text is mine; the English translation, the investigation, the patch, the test and the measurements were done with AI assistance (Claude). The tables are the measured numbers.

@midagedev
midagedev requested review from a team and pwilkin as code owners September 17, 2026 02:35
@github-actions github-actions Bot added the testing Everything test related label Sep 17, 2026
@aldehir

aldehir commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

Please remove the tests and we can get this merged in.

@midagedev
midagedev force-pushed the deepseek-msg-delimiters branch from b861638 to a86f622 Compare September 17, 2026 03:14
@midagedev

Copy link
Copy Markdown
Contributor Author

Please remove the tests and we can get this merged in.

Done, removed the test. thanks

@aldehir aldehir added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 17, 2026
@CISC

CISC commented Sep 17, 2026

Copy link
Copy Markdown
Member

The attempt to align something that won't align was just confusing. :)

@pwilkin
pwilkin merged commit 7f6f0c2 into ggml-org:master Sep 17, 2026
25 of 26 checks passed
CISC added a commit that referenced this pull request Sep 17, 2026
* chat : add message delimiters to the DeepSeek V3.2/V4 parser

Assisted-by: Claude
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
midagedev added a commit to midagedev/rig-log that referenced this pull request Sep 17, 2026
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@NovNovikov

Copy link
Copy Markdown

The bug that you fixed actually protected deepseek from the questionable solution about checkpoints on every user turn that ruins prefill speed completely on long sessions. Now try to restart llama and recalculate your whole prefil and compare the prompt processing speed. #25320

fencerJP pushed a commit to fencerJP/llama-apu that referenced this pull request Sep 19, 2026
…g#29008)

* chat : add message delimiters to the DeepSeek V3.2/V4 parser

Assisted-by: Claude
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
@midagedev

Copy link
Copy Markdown
Contributor Author

@NovNovikov You are right that I never tested that shape — the table in this PR is a
huge system prompt with one user turn. So I measured it.

No rebuild needed: the delimiters travel in the request, so on one server with one set
of tokens, ON = /v1/chat/completions, OFF = /completion with the token array from
/apply-template + /tokenize (no message_delimiters field = pre-patch behaviour).
prompt_n was identical in every pair.

DeepSeek-V4-Flash-0731 UD-Q4_K_XL, -b 2048 -ub 2048 -np 1 -ngl 99, experts mostly on
host, cache_prompt: false, idle box. Cold prefill, ms, chunk count in parentheses:

--checkpoint-min-step 48160 tok, 121 user turns 17522 tok, 1 user turn
ON OFF ON OFF
0 438120 (123) 134450 (24) 53805 (10) 53616 (9)
8192 (default) 154972 (31) 136633 (24) 54852 (10) 53901 (9)
65536 141803 (27) 135875 (24) 53935 (10) 54069 (9)

ON/OFF: 3.26× / 1.13× / 1.04× on the long chat; 1.00–1.02× on the short one.

min-step 0 is #25320 — 121 turns become 123 chunks, the smallest 5 tokens. But that
issue was closed on 2026-07-09 by the grouping guard, and the default has been 8192
since before your comment, which holds the same conversation to 1.13× (31 chunks:
~5.9 from 48160/8192, plus the first user start, the last one, and the two tail
checkpoints). At 65536 only those four remain and it is 1.04×.

So the cost is real on long many-turn sessions, it is bounded by a knob that already
exists, and it is zero on the shape this PR was about. Happy to open a separate issue
if you think 13% at defaults is worth changing the break rule — I did not find an open
one after #25320.

LadislavSopko pushed a commit to 0ics-srls/llama.cpp that referenced this pull request Oct 5, 2026
…g#29008)

* chat : add message delimiters to the DeepSeek V3.2/V4 parser

Assisted-by: Claude
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…g#29008)

* chat : add message delimiters to the DeepSeek V3.2/V4 parser

Assisted-by: Claude
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants