Skip to content

ggml : fix ggml_backend_buft_get_alloc_size() guard - #28038

Merged
ggerganov merged 1 commit into
masterfrom
gg/ggml-may-expand-fix
Aug 30, 2026
Merged

ggerganov merged 1 commit into
masterfrom
gg/ggml-may-expand-fix

Conversation

@ggerganov

Copy link
Copy Markdown
Member

Overview

cont #27960

Didn't take into account that CUDA pads quantized tensors.

Fix: https://github.com/ggml-org/llama.cpp/actions/runs/33296587275/job/99217196889#step:3:31473

Requirements

@ggerganov
ggerganov requested a review from a team as a code owner August 30, 2026 16:42
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Aug 30, 2026
@ggerganov
ggerganov merged commit 6d1479c into master Aug 30, 2026
26 of 30 checks passed
@ggerganov
ggerganov deleted the gg/ggml-may-expand-fix branch August 30, 2026 17:25
thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
SteelPh0enix pushed a commit to SteelPh0enix/llama.cpp-qwen4exp that referenced this pull request Sep 8, 2026
SteelPh0enix pushed a commit to SteelPh0enix/llama.cpp-qwen4exp that referenced this pull request Sep 8, 2026
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
Jesssullivan added a commit to Jesssullivan/llama.cpp that referenced this pull request Sep 20, 2026
PLACEHOLDER MESSAGE — rewrite before opening a pull request. llama.cpp
CONTRIBUTING prohibits AI-written commit messages and PR descriptions;
this text was prepared by tooling so the branch could be pushed, and the
author is expected to replace it with their own words (git commit --amend).

ggml_backend_buft_get_alloc_size guards the documented lower bound of the
buffer-type contract with a plain assert(), which is compiled out under
NDEBUG, while the upper bound two lines below is a GGML_ASSERT. A buffer
type that under-reports therefore passes silently in release builds and
only fails later, in the in-place-reuse path, at
ggml-alloc.c GGML_ASSERT(parent_size >= node_size).

Promote the lower bound to GGML_ASSERT so the failure lands on the
buffer type that produced the wrong size. The neighbouring assert already
evaluates ggml_nbytes in release builds, so there is no added cost.

ref: ggml-org#27960
ref: ggml-org#28038
ref: ggml-org#17884

Assisted-by: Claude Opus 5
Jesssullivan added a commit to Jesssullivan/llama.cpp that referenced this pull request Sep 20, 2026
… at ggml_nbytes

PLACEHOLDER MESSAGE — rewrite before opening a pull request. llama.cpp
CONTRIBUTING prohibits AI-written commit messages and PR descriptions;
this text was prepared by tooling so the branch could be pushed, and the
author is expected to replace it with their own words (git commit --amend).

The process-global memo added in ggml-org#18626 keys RPC_CMD_GET_ALLOC_SIZE
results on {device, type, op, op_params, ne} and omits nb. Two tensors
with identical ne but different strides therefore share one entry, and
the second caller receives the first tensor's size. When that size is
smaller than ggml_nbytes, the buffer-type contract documented in ggml-org#27960
and ggml-org#28038 is violated and allocation fails downstream.

Add nb to the cache key and never return less than ggml_nbytes. The memo
is preserved: contiguous tensors, which are the common case, keep the
same hit rate. rpc_tensor already carries nb over the wire, so there is
no protocol or API change and RPC_PROTO_* is untouched.

Reproduced against a local ggml-rpc-server -d CPU: a packed tensor
reports 16384/16384 and a strided view with the same ne reports
nbytes 32512, alloc_size 16384, then aborts; with this change the second
call reports 32512 and the run completes. The regression is folded into
tests/test-rpc-multi-server.

ref: ggml-org#28360
ref: ggml-org#18626
ref: ggml-org#27960

Assisted-by: Claude Opus 5
Jesssullivan added a commit to Jesssullivan/llama.cpp that referenced this pull request Sep 22, 2026
PLACEHOLDER MESSAGE — rewrite before opening a pull request. llama.cpp
CONTRIBUTING prohibits AI-written commit messages and PR descriptions;
this text was prepared by tooling so the branch could be pushed, and the
author is expected to replace it with their own words (git commit --amend).

ggml_backend_buft_get_alloc_size guards the documented lower bound of the
buffer-type contract with a plain assert(), which is compiled out under
NDEBUG, while the upper bound two lines below is a GGML_ASSERT. A buffer
type that under-reports therefore passes silently in release builds and
only fails later, in the in-place-reuse path, at
ggml-alloc.c GGML_ASSERT(parent_size >= node_size).

Promote the lower bound to GGML_ASSERT so the failure lands on the
buffer type that produced the wrong size. The neighbouring assert already
evaluates ggml_nbytes in release builds, so there is no added cost.

ref: ggml-org#27960
ref: ggml-org#28038
ref: ggml-org#17884

Assisted-by: Claude Opus 5
Jesssullivan added a commit to Jesssullivan/llama.cpp that referenced this pull request Sep 22, 2026
… at ggml_nbytes

PLACEHOLDER MESSAGE — rewrite before opening a pull request. llama.cpp
CONTRIBUTING prohibits AI-written commit messages and PR descriptions;
this text was prepared by tooling so the branch could be pushed, and the
author is expected to replace it with their own words (git commit --amend).

The process-global memo added in ggml-org#18626 keys RPC_CMD_GET_ALLOC_SIZE
results on {device, type, op, op_params, ne} and omits nb. Two tensors
with identical ne but different strides therefore share one entry, and
the second caller receives the first tensor's size. When that size is
smaller than ggml_nbytes, the buffer-type contract documented in ggml-org#27960
and ggml-org#28038 is violated and allocation fails downstream.

Add nb to the cache key and never return less than ggml_nbytes. The memo
is preserved: contiguous tensors, which are the common case, keep the
same hit rate. rpc_tensor already carries nb over the wire, so there is
no protocol or API change and RPC_PROTO_* is untouched.

Reproduced against a local ggml-rpc-server -d CPU: a packed tensor
reports 16384/16384 and a strided view with the same ne reports
nbytes 32512, alloc_size 16384, then aborts; with this change the second
call reports 32512 and the run completes. The regression is folded into
tests/test-rpc-multi-server.

ref: ggml-org#28360
ref: ggml-org#18626
ref: ggml-org#27960

Assisted-by: Claude Opus 5
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant