Skip to content

server : share checkpoint state and harden prompt cache OOM handling - #27451

Open
Hundsbuah wants to merge 1 commit into
ggml-org:masterfrom
Hundsbuah:fix/shared-prompt-checkpoints
Open

Hundsbuah wants to merge 1 commit into
ggml-org:masterfrom
Hundsbuah:fix/shared-prompt-checkpoints

Conversation

@Hundsbuah

@Hundsbuah Hundsbuah commented Aug 20, 2026 •

Copy link
Copy Markdown

Share target and draft checkpoint backing storage to avoid large deep copies when saving prompts to the in-memory cache.

Also handle std::bad_alloc during complete cache entry construction, including states.push_back().

Assisted-by: Qwen3.8-27B

Summary

This PR prevents excessive transient host-memory usage when llama-server
saves prompts containing large context checkpoints into the in-memory
prompt cache.

The current implementation deep-copies the complete checkpoint payload
when a prompt is inserted into the cache. For long-context
hybrid/recurrent workloads, checkpoint state can become large enough for
this temporary duplication to cause significant memory pressure or
allocation failure even when the resulting cache entry itself satisfies
the configured --cache-ram limit.

This PR:

  • shares the large target and draft checkpoint state buffers using
    reference-counted backing storage;
  • preserves checkpoint metadata and rollback semantics;
  • keeps checkpoints available for hybrid/recurrent prompt reuse;
  • avoids deep-copying checkpoint payloads during server_prompt copies;
  • extends std::bad_alloc handling to cover complete prompt-cache entry
    construction, including states.push_back(...).

Problem

server_prompt stores checkpoints by value:

struct server_prompt {
    server_tokens tokens;
    std::list<common_prompt_checkpoint> checkpoints;
};

Each checkpoint contains serialized state buffers such as:

std::vector<uint8_t> data_tgt;
std::vector<uint8_t> data_dft;

When a prompt is saved to the in-memory prompt cache,
server_prompt_cache::alloc() constructs a new cache entry using the
equivalent of:

states.push_back({
    {
        prompt.tokens.clone(),
        prompt.checkpoints,
    },
    ...
});

Copying prompt.checkpoints therefore performs a deep copy of every
data_tgt and data_dft buffer.

For long contexts or models requiring substantial recurrent checkpoint
state, this can make prompt-cache insertion temporarily require close to
an additional copy of the complete checkpoint payload.

Why --cache-ram does not prevent the transient allocation

The cache admission logic includes checkpoint sizes when calculating the
logical size of a new cache entry:

size_t checkpoints_size = 0;

for (const auto & ckpt : prompt.checkpoints) {
    checkpoints_size += ckpt.size();
}

const size_t state_size_new =
    state_size_tgt +
    state_size_dft +
    checkpoints_size;

This correctly limits the logical residency of the resulting cache entry.

However, during cache-entry construction the active prompt still owns its
original checkpoint buffers while the cache creates another copy.

The transient memory requirement is therefore approximately:

existing checkpoint payload
+ newly allocated target/draft prompt state
+ duplicated checkpoint payload

As a result, satisfying the logical cache-size limit does not guarantee
that enough host memory is available for the temporary deep copy required
to construct the cache entry.

Why checkpoints should not simply be removed

Clearing checkpoints before saving the prompt would avoid the duplication,
but would change behavior for hybrid/recurrent models.

These models may require captured checkpoints to restore recurrent state
when a later prompt shares only part of a cached prefix.

Discarding those checkpoints can prevent efficient partial-prefix reuse
and require substantially more prompt reprocessing.

This PR therefore preserves the checkpoints and removes only the
unnecessary duplication of their large serialized backing storage.

Implementation

Shared checkpoint payload

The large serialized target and draft checkpoint buffers now use
reference-counted backing storage.

Conceptually, instead of:

active prompt ───> checkpoint payload A

cached prompt ───> duplicated checkpoint payload A

copies now behave as:

active prompt ─┐
               ├──> checkpoint payload A
cached prompt ─┘

Only the lightweight ownership handle is copied.

Snapshot semantics are preserved

Checkpoint payloads are treated as immutable after capture.

When a checkpoint is updated, a new backing buffer is allocated and the
new serialized state is written into that buffer rather than modifying an
existing shared snapshot.

Conceptually:

before update:

slot  ─┐
       ├──> snapshot A
cache ─┘

after slot update:

cache ───> snapshot A
slot  ───> snapshot B

This preserves independent checkpoint snapshot semantics while avoiding
unnecessary deep copies.

Target and draft state

Both large serialized state buffers use shared backing storage:

data_tgt
data_dft

This preserves the behavior required for target and draft model state,
including speculative/MTP workloads.

The smaller speculative checkpoint state remains value-owned to keep the
change focused and avoid unnecessary modifications outside the primary
memory-pressure path.

OOM handling

The existing implementation catches std::bad_alloc around:

state_data_tgt.resize(state_size_tgt);
state_data_dft.resize(state_size_dft);

However, subsequent operations may allocate memory as well:

prompt.tokens.clone()
prompt.checkpoints
states.push_back(...)

These operations were previously outside the protected section.

This PR extends the try block to cover the complete cache-entry
construction:

try {
    state_data_tgt.resize(state_size_tgt);
    state_data_dft.resize(state_size_dft);

    states.push_back({
        {
            prompt.tokens.clone(),
            prompt.checkpoints,
        },
        ...
    });
} catch (const std::bad_alloc & e) {
    ...
}

If any allocation involved in constructing the cache entry fails, the
existing cache-recovery path is used and the function returns without
leaving a partially constructed entry.

Temporary allocations are released automatically during stack unwinding.

Cache accounting

This PR intentionally does not change the existing logical cache-size
accounting.

Checkpoint payload sizes continue to count toward --cache-ram even when
their physical backing storage is shared.

This can make cache accounting conservative when multiple prompts refer
to the same backing allocation, potentially causing earlier eviction than
strict physical-byte accounting would require.

Keeping the existing accounting is intentional because it:

  • preserves current --cache-ram semantics;
  • avoids introducing under-counting;
  • keeps the fix focused on eliminating unnecessary duplication;
  • avoids coupling this change to a broader cache-accounting redesign.

Unique physical backing-store accounting or a global checkpoint pool can
be considered separately.

Expected memory behavior

Before this change, prompt-cache insertion may temporarily require:

checkpoint payload
+ target/draft prompt state
+ second copy of checkpoint payload

With shared checkpoint backing storage, the same operation requires
approximately:

checkpoint payload
+ target/draft prompt state
+ lightweight checkpoint references

The large checkpoint payload is therefore retained only once instead of
being duplicated during cache insertion.

This does not make --cache-ram a total process-memory limit. Overall
memory usage still depends on active slots, model state, compute buffers,
allocator behavior and other process allocations.

Scope

The patch is deliberately limited to:

  • checkpoint backing-storage ownership;
  • checkpoint update and clear handling required for shared snapshots;
  • prompt-cache allocation exception handling.

It does not change:

  • checkpoint selection;
  • checkpoint creation policy;
  • rollback logic;
  • prompt-cache matching;
  • cache eviction policy;
  • speculative decoding behavior;
  • recurrent-state algorithms;
  • logical --cache-ram accounting.

Validation

The patch has been applied and checked against the target source tree.

The affected translation units compile successfully:

common/common.cpp
tools/server/server-task.cpp
tools/server/server-context.cpp

The recurrent-state rollback test build also proceeds through the affected
code without errors.

Additional runtime validation is appropriate for:

  • transformer-only models;
  • hybrid/recurrent models;
  • prompt-cache hit and miss paths;
  • partial-prefix reuse after cache restore;
  • repeated checkpoint replacement;
  • cache eviction while checkpoint backing storage is shared;
  • speculative/MTP target and draft checkpoint restoration;
  • large-context workloads;
  • forced allocation failure to verify graceful std::bad_alloc handling;
  • ownership and lifetime testing under sanitizers where available.

Motivation

The goal is to eliminate excessive transient checkpoint duplication
without sacrificing the state required for correct and efficient
hybrid/recurrent prompt reuse.

The core ownership model becomes:

checkpoint metadata:
    value semantics

serialized checkpoint payload:
    immutable shared backing storage

checkpoint update:
    allocate a new snapshot

checkpoint destruction / cache eviction:
    release one reference

prompt-cache copy:
    copy references rather than serialized state

This removes the problematic deep copy while retaining checkpoint
information needed for recurrent-state restoration and prompt reuse.

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES - AI was used for bug finding and fixing

Share target and draft checkpoint backing storage to avoid large deep copies when saving prompts to the in-memory cache.

Also handle `std::bad_alloc` during complete cache entry construction, including `states.push_back()`.

Assisted-by: Qwen3.8-27B
@Hundsbuah
Hundsbuah requested review from a team as code owners August 20, 2026 17:29
@ggml-gh-bot

ggml-gh-bot Bot commented Aug 20, 2026

Copy link
Copy Markdown

Hi @Hundsbuah, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@ggml-gh-bot ggml-gh-bot Bot added the draft PR will be changed to draft by github-actions bot label Aug 20, 2026
@github-actions
github-actions Bot marked this pull request as draft August 20, 2026 17:35
@github-actions github-actions Bot removed the draft PR will be changed to draft by github-actions bot label Aug 20, 2026
@Hundsbuah
Hundsbuah marked this pull request as ready for review August 20, 2026 17:47
benjigill added a commit to benjigill/llama.cpp that referenced this pull request Sep 23, 2026
…cal)

bench_qwen38.py cache suite, branch vs master build:

                        master   branch
  cold ttft s            19.72    19.89   +0.9%
  warm prompt_n       18/16/17 18/16/17
  multiturn prompt_n     26/22    26/22
  multiturn ttft s   0.52/0.52 0.54/0.52

No regression; the suite does not reach the long-prefill lockup or the
--cache-ram limit these patches fix.

Assisted-by: Claude Opus 5.5
@danilofalcao

Copy link
Copy Markdown

Related behavior we hit on a recurrent hybrid served over RPC: when the draft KV is full in the middle of a speculative pass, we purge idle slots and retry the draft pass once; if that still fails we clear the out-of-sync slots instead of failing the request. Under 3 concurrent sessions on a 64K unified pool this stopped speculative decoding from degrading the whole session on KV exhaustion.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants