Repository navigation
chat : add message delimiters to the DeepSeek V3.2/V4 parser - #29008
Conversation
|
Please remove the tests and we can get this merged in. |
Assisted-by: Claude
b861638 to
a86f622
Compare
Done, removed the test. thanks |
|
The attempt to align something that won't align was just confusing. :) |
* chat : add message delimiters to the DeepSeek V3.2/V4 parser Assisted-by: Claude Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
The bug that you fixed actually protected deepseek from the questionable solution about checkpoints on every user turn that ruins prefill speed completely on long sessions. Now try to restart llama and recalculate your whole prefil and compare the prompt processing speed. #25320 |
…g#29008) * chat : add message delimiters to the DeepSeek V3.2/V4 parser Assisted-by: Claude Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
|
@NovNovikov You are right that I never tested that shape — the table in this PR is a No rebuild needed: the delimiters travel in the request, so on one server with one set DeepSeek-V4-Flash-0731 UD-Q4_K_XL,
ON/OFF: 3.26× / 1.13× / 1.04× on the long chat; 1.00–1.02× on the short one.
So the cost is real on long many-turn sessions, it is bounded by a knob that already |
…g#29008) * chat : add message delimiters to the DeepSeek V3.2/V4 parser Assisted-by: Claude Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
…g#29008) * chat : add message delimiters to the DeepSeek V3.2/V4 parser Assisted-by: Claude Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
Assisted-by: Claude
Overview
안녕하세요. 영어를 잘 하지 못해서 한국어로 쓰고, 아래에 번역을 붙입니다.
DeepSeek V4 Flash를 코딩 에이전트에서 쓰다가 프리필이 지나치게 오래 걸리는 문제를 만났습니다. 분석해 보니 다른 모델과 달리 DeepSeek 전용 파서가 유저 메시지 구분자를 서버에 알려주지 않아서, 유저 턴 위치에 체크포인트가 생기지 않고 프롬프트 캐시가 거의 활용되지 않는 것이 원인이었습니다.
포크해서 파서에 구분자 두 개를 추가했더니 캐시 적중이 명확하게 올랐습니다. #24176이 같은 처리를 다른 파서들에 넣었는데, 그 시점에 이미 있던 DeepSeek V3.2 파서는 목록에 빠져 있었습니다.
검증 자료를 아래에 첨부했으니 검토 부탁드립니다.
(English, translated from the Korean above)
Hello. My English is not good, so I wrote this in Korean; a translation follows.
While using DeepSeek V4 Flash with a coding agent, prefill took far too long. The cause turned out to be different from other models: the DeepSeek specialized parser does not report the user message delimiters to the server, so no checkpoint is created at user turns and the prompt cache is barely used.
I added the two delimiters to the parser in a fork and cache reuse went up clearly. #24176 added the same thing to the other parsers, but the DeepSeek V3.2 parser, which already existed at that time, was not in that list.
Verification data is attached below.
Change:
common_chat_params_init_deepseek_v3_2now setsdata.message_delimitersto USER<|User|>and ASSISTANT<|Assistant|>. Both markers open every turn in the V3.2, V4 and V4.1 templates.Test:
test_deepseek_message_delimitersintests/test-chat.cpp, V3.2 and V4 templates. Fails on master (Expected: <|User|> Actual:empty, rc 134), passes with the patch. Same style astest_msg_token_delimiters_splitnext to it; no new test file.Additional information
Why it matters for this model: DeepSeek V3.2/V4/V4.1 use SWA layers (
n_swa128,llama_kv_cache_dsv4). The KV cache cannot be rolled back to a position without a checkpoint, and without delimiters the server only places checkpoints atend - 4andend - 4 - n_ubatch. So any request that diverges before the tail re-prefills from token 0.Setup: DeepSeek-V4.1-Flash, one RTX A6000, experts on host (
-ot "exps=CPU"), llama-server default checkpoint settings. Prompt = first request of a coding agent, 13 167 tokens (4 200 system, ~8 950 tool schemas for 23 tools, 20 user). Prefill ~46 tok/s. Numbers fromtimingsin the response.master (checkpoints in the log:
n_tokens12 651 and 13 163 only):cache_nprompt_n--ctx-checkpoints 64 --checkpoint-min-step 1024patched (third checkpoint at
n_tokens13 145 = first user message):cache_nprompt_nNot tested:
Unchanged by this PR: an edit inside the system prompt or the tool list still re-prefills from 0. That is the mid-prompt checkpoint policy from #22929, not the parser.
Same omission, not touched here:
common/parsers/lfm2.cpp(LFM2 is recurrent + SWA, so it should behave the same; user marker would be<|im_start|>user\n).ministral3,gigachat_v3,kimi_k2,functionary_v3_2also lack delimiters but serve full-KV models where no checkpoint is needed.Related: #24176, #21785 (the V3.2 parser), #25452 (DSV4-Flash divergent turns re-prefill from the checkpoint boundary), #21831, #24055, #28302.
Searched before opening:
gh pr list --search"message_delimiters", "deepseek delimiters", "deepseek checkpoint";gh issue list --search"deepseek prompt cache", "V4 checkpoint", "cache_n deepseek"; open PRs touchingcommon/parsers/deepseek.cpp: #28612, #28724 (unrelated).Requirements