Repository navigation
models: pad on the left with ggml_pad_ext - #29567
Merged
ServeurpersoCom merged 2 commits intoSep 28, 2026
Merged
Conversation
The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders build a left padding as a right pad followed by a roll, and DFlash2 concatenates a zero filled block in front of the previous tokens. ggml_pad_ext does both in one node now that every backend supports a left padding. The Gemma 4 audio embeddings are bit identical.
ggerganov
approved these changes
Sep 28, 2026
ggerganov
reviewed
Sep 28, 2026
| values = ggml_fill(ctx0, | ||
| ggml_new_tensor_3d(ctx0, hidden->type, hidden_size, block_size, n_blocks), 0.0f); | ||
| } | ||
| } |
Member
There was a problem hiding this comment.
This section might be possible to simplify:
if (tap > 0) {
tap_cur = std::min(tap, block_size);
ggml_tensor * previous = ggml_view_3d(ctx0, blocks, hidden_size, block_size - tap_cur, n_blocks,
blocks->nb[1], blocks->nb[2], 0);
values = ggml_pad_ext(ctx0, previous, 0, 0, tap_cur, 0, 0, 0, 0, 0);
}
Contributor
Author
There was a problem hiding this comment.
Thanks, went one step further: a tap at or past block_size only reads the padding, so its term is zero and it is skipped, the else branch goes away. Tested, the result is mathematically identical.
A tap at or past block_size shifts every row out of the block, so its term is zero. The loop runs min(kernel_size, block_size) taps.
Contributor
Author
|
Tested bit-exact against master on CPU, CUDA, Vulkan and Metal with a standalone ggml test: it builds the master, previous and current versions of the tap loop on random data, runs them on every backend and compares the outputs with memcmp, over block sizes above, equal to and below the kernel size. ggml test// Compares three builds of the DFlash2 tap loop on random data, bit for bit:
// master (zero block concatenated in front), PR (pad_ext plus fill for dead taps)
// and the reviewed version (pad_ext, dead taps skipped).
#include "ggml.h"
#include "ggml-alloc.h"
#include "ggml-backend.h"
#include "ggml-cpu.h"
#include <algorithm>
#include <cstdio>
#include <cstring>
#include <random>
#include <vector>
enum { V_MASTER, V_PR, V_NEW };
static ggml_tensor * build(ggml_context * ctx0, ggml_tensor * hidden, ggml_tensor * weight_all,
int64_t kernel_size, int64_t n_blocks, int variant) {
const int64_t hidden_size = hidden->ne[0];
const int64_t n_tokens = hidden->ne[1];
const int64_t block_size = n_tokens / n_blocks;
const int64_t group_size = weight_all->ne[0];
const int64_t n_groups = weight_all->ne[1];
ggml_tensor * blocks = ggml_reshape_3d(ctx0, hidden, hidden_size, block_size, n_blocks);
const int64_t n_taps = variant == V_NEW ? std::min(kernel_size, block_size) : kernel_size;
ggml_tensor * result = nullptr;
for (int64_t tap = 0; tap < n_taps; ++tap) {
ggml_tensor * values = blocks;
if (tap > 0) {
if (variant == V_MASTER) {
ggml_tensor * zeros = ggml_fill(ctx0,
ggml_new_tensor_3d(ctx0, hidden->type, hidden_size, std::min(tap, block_size), n_blocks), 0.0f);
if (tap < block_size) {
ggml_tensor * previous = ggml_view_3d(ctx0, blocks, hidden_size, block_size - tap, n_blocks,
blocks->nb[1], blocks->nb[2], 0);
values = ggml_concat(ctx0, zeros, previous, 1);
} else {
values = zeros;
}
} else if (variant == V_PR && tap >= block_size) {
values = ggml_fill(ctx0,
ggml_new_tensor_3d(ctx0, hidden->type, hidden_size, block_size, n_blocks), 0.0f);
} else {
ggml_tensor * previous = ggml_view_3d(ctx0, blocks, hidden_size, block_size - tap, n_blocks,
blocks->nb[1], blocks->nb[2], 0);
values = ggml_pad_ext(ctx0, previous, 0, 0, tap, 0, 0, 0, 0, 0);
}
}
values = ggml_reshape_2d(ctx0, values, hidden_size, n_tokens);
ggml_tensor * weight = ggml_reshape_2d(ctx0,
ggml_cont(ctx0, ggml_view_4d(ctx0, weight_all, group_size, n_groups, 1, n_tokens,
weight_all->nb[1], weight_all->nb[2], weight_all->nb[3], tap * weight_all->nb[2])),
hidden_size, n_tokens);
ggml_tensor * term = ggml_mul(ctx0, weight, values);
result = result ? ggml_add(ctx0, result, term) : term;
}
return result;
}
static std::vector<float> run(ggml_backend_t be, const std::vector<float> & h, const std::vector<float> & w,
int64_t H, int64_t T, int64_t G, int64_t K, int64_t NB, int variant, int * n_nodes) {
ggml_init_params ip = { 64*1024*1024, nullptr, true };
ggml_context * ctx = ggml_init(ip);
ggml_tensor * hidden = ggml_new_tensor_2d(ctx, GGML_TYPE_F32, H, T);
ggml_tensor * wall = ggml_new_tensor_4d(ctx, GGML_TYPE_F32, G, H / G, K, T);
ggml_tensor * out = build(ctx, hidden, wall, K, NB, variant);
ggml_cgraph * gf = ggml_new_graph(ctx);
ggml_build_forward_expand(gf, out);
*n_nodes = ggml_graph_n_nodes(gf);
ggml_backend_buffer_t buf = ggml_backend_alloc_ctx_tensors(ctx, be);
ggml_backend_tensor_set(hidden, h.data(), 0, h.size()*4);
ggml_backend_tensor_set(wall, w.data(), 0, w.size()*4);
ggml_backend_graph_compute(be, gf);
std::vector<float> r(ggml_nelements(out));
ggml_backend_tensor_get(out, r.data(), 0, r.size()*4);
ggml_backend_buffer_free(buf);
ggml_free(ctx);
return r;
}
int main() {
ggml_backend_load_all();
std::vector<ggml_backend_t> bes;
for (size_t i = 0; i < ggml_backend_dev_count(); ++i) {
ggml_backend_dev_t d = ggml_backend_dev_get(i);
if (ggml_backend_dev_type(d) != GGML_BACKEND_DEVICE_TYPE_ACCEL) {
bes.push_back(ggml_backend_dev_init(d, nullptr));
}
}
// {hidden, group, kernel, block_size, n_blocks}
const int64_t cases[][5] = {
{64, 16, 4, 16, 3}, {64, 16, 4, 4, 2}, {64, 16, 4, 3, 5}, {64, 16, 4, 1, 7},
{128, 32, 7, 2, 4}, {256, 64, 3, 16, 1}, {32, 8, 5, 5, 1},
};
std::mt19937 rng(42);
std::normal_distribution<float> nd;
int fails = 0;
for (auto & c : cases) {
const int64_t H = c[0], G = c[1], K = c[2], BS = c[3], NB = c[4], T = BS*NB;
std::vector<float> h(H*T), w(G*(H/G)*K*T);
for (auto & x : h) x = nd(rng);
for (auto & x : w) x = nd(rng);
for (auto be : bes) {
int nm, np, nn;
auto rm = run(be, h, w, H, T, G, K, NB, V_MASTER, &nm);
auto rp = run(be, h, w, H, T, G, K, NB, V_PR, &np);
auto rn = run(be, h, w, H, T, G, K, NB, V_NEW, &nn);
const bool okp = memcmp(rm.data(), rp.data(), rm.size()*4) == 0;
const bool okn = memcmp(rm.data(), rn.data(), rm.size()*4) == 0;
fails += !okp + !okn;
printf("%-6s H=%-3lld K=%lld block=%-2lld blocks=%lld nodes master/pr/new %2d/%2d/%2d pr %s new %s\n",
ggml_backend_name(be), (long long)H, (long long)K, (long long)BS, (long long)NB, nm, np, nn,
okp ? "bit-exact" : "DIFF", okn ? "bit-exact" : "DIFF");
}
}
printf("%s\n", fails ? "FAIL" : "ALL BIT-EXACT");
for (auto be : bes) ggml_backend_free(be);
return fails != 0;
} |
ggerganov
approved these changes
Sep 28, 2026
ngxson
approved these changes
Sep 28, 2026
pierreguillot
pushed a commit
to Ircam-Partiels/llama.cpp
that referenced
this pull request
Oct 1, 2026
* models: pad on the left with ggml_pad_ext The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders build a left padding as a right pad followed by a roll, and DFlash2 concatenates a zero filled block in front of the previous tokens. ggml_pad_ext does both in one node now that every backend supports a left padding. The Gemma 4 audio embeddings are bit identical. * models: skip the DFlash2 taps that only read padding A tap at or past block_size shifts every row out of the block, so its term is zero. The loop runs min(kernel_size, block_size) taps.
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
* models: pad on the left with ggml_pad_ext The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders build a left padding as a right pad followed by a roll, and DFlash2 concatenates a zero filled block in front of the previous tokens. ggml_pad_ext does both in one node now that every backend supports a left padding. The Gemma 4 audio embeddings are bit identical. * models: skip the DFlash2 taps that only read padding A tap at or past block_size shifts every row out of the block, so its term is zero. The loop runs min(kernel_size, block_size) taps.
edwardyoon
pushed a commit
to edwardyoon/focus-llama
that referenced
this pull request
Oct 7, 2026
* models: pad on the left with ggml_pad_ext The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders build a left padding as a right pad followed by a roll, and DFlash2 concatenates a zero filled block in front of the previous tokens. ggml_pad_ext does both in one node now that every backend supports a left padding. The Gemma 4 audio embeddings are bit identical. * models: skip the DFlash2 taps that only read padding A tap at or past block_size shifts every row out of the block, so its term is zero. The loop runs min(kernel_size, block_size) taps. (cherry picked from commit 57b557c)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Several models build a left padding as a right pad followed by a roll (Parakeet, LFM2-Audio, Granite Speech, Gemma 4 audio), or as a zero block concatenated in front of the data (DFlash2). This replaces each of them with a single ggml_pad_ext.
The result is mathematically identical, since the roll only moves the padded zeros, and each site saves one node and one memory pass. Gemma 4 E2B audio embeddings are bit identical before and after.
Additional information
Follow-up to #29561, which made left padding available on every backend including Metal.
Requirements