Skip to content

models: pad on the left with ggml_pad_ext - #29567

Merged
ServeurpersoCom merged 2 commits into
ggml-org:masterfrom
ServeurpersoCom:models-pad-ext-left
Sep 28, 2026
Merged

ServeurpersoCom merged 2 commits into
ggml-org:masterfrom
ServeurpersoCom:models-pad-ext-left

Conversation

@ServeurpersoCom

@ServeurpersoCom ServeurpersoCom commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Several models build a left padding as a right pad followed by a roll (Parakeet, LFM2-Audio, Granite Speech, Gemma 4 audio), or as a zero block concatenated in front of the data (DFlash2). This replaces each of them with a single ggml_pad_ext.

The result is mathematically identical, since the roll only moves the padded zeros, and each site saves one node and one memory pass. Gemma 4 E2B audio embeddings are bit identical before and after.

Additional information

Follow-up to #29561, which made left padding available on every backend including Metal.

Requirements

The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders
build a left padding as a right pad followed by a roll, and DFlash2
concatenates a zero filled block in front of the previous tokens.
ggml_pad_ext does both in one node now that every backend supports a
left padding. The Gemma 4 audio embeddings are bit identical.
@ServeurpersoCom
ServeurpersoCom requested review from a team and CISC as code owners September 28, 2026 11:11
@github-actions github-actions Bot added model Model specific mtmd Related to multimodal functionality (video/image/audio) labels Sep 28, 2026
Comment thread src/models/dflash.cpp
values = ggml_fill(ctx0,
ggml_new_tensor_3d(ctx0, hidden->type, hidden_size, block_size, n_blocks), 0.0f);
}
}

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This section might be possible to simplify:

        if (tap > 0) {
            tap_cur = std::min(tap, block_size);
            ggml_tensor * previous = ggml_view_3d(ctx0, blocks, hidden_size, block_size - tap_cur, n_blocks,
                        blocks->nb[1], blocks->nb[2], 0);
            values = ggml_pad_ext(ctx0, previous, 0, 0, tap_cur, 0, 0, 0, 0, 0);
        }

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, went one step further: a tap at or past block_size only reads the padding, so its term is zero and it is skipped, the else branch goes away. Tested, the result is mathematically identical.

A tap at or past block_size shifts every row out of the block, so its
term is zero. The loop runs min(kernel_size, block_size) taps.
@ServeurpersoCom

Copy link
Copy Markdown
Contributor Author

Tested bit-exact against master on CPU, CUDA, Vulkan and Metal with a standalone ggml test: it builds the master, previous and current versions of the tap loop on random data, runs them on every backend and compares the outputs with memcmp, over block sizes above, equal to and below the kernel size.

ggml test
c++ -O2 -std=c++17 dflash_check.cpp -Iggml/include -Lbuild/bin -lggml -lggml-base -lggml-cpu -Wl,-rpath,$PWD/build/bin -o dflash_check && ./dflash_check
// Compares three builds of the DFlash2 tap loop on random data, bit for bit:
// master (zero block concatenated in front), PR (pad_ext plus fill for dead taps)
// and the reviewed version (pad_ext, dead taps skipped).
#include "ggml.h"
#include "ggml-alloc.h"
#include "ggml-backend.h"
#include "ggml-cpu.h"
#include <algorithm>
#include <cstdio>
#include <cstring>
#include <random>
#include <vector>

enum { V_MASTER, V_PR, V_NEW };

static ggml_tensor * build(ggml_context * ctx0, ggml_tensor * hidden, ggml_tensor * weight_all,
        int64_t kernel_size, int64_t n_blocks, int variant) {
    const int64_t hidden_size = hidden->ne[0];
    const int64_t n_tokens    = hidden->ne[1];
    const int64_t block_size  = n_tokens / n_blocks;
    const int64_t group_size  = weight_all->ne[0];
    const int64_t n_groups    = weight_all->ne[1];
    ggml_tensor * blocks = ggml_reshape_3d(ctx0, hidden, hidden_size, block_size, n_blocks);

    const int64_t n_taps = variant == V_NEW ? std::min(kernel_size, block_size) : kernel_size;
    ggml_tensor * result = nullptr;
    for (int64_t tap = 0; tap < n_taps; ++tap) {
        ggml_tensor * values = blocks;
        if (tap > 0) {
            if (variant == V_MASTER) {
                ggml_tensor * zeros = ggml_fill(ctx0,
                        ggml_new_tensor_3d(ctx0, hidden->type, hidden_size, std::min(tap, block_size), n_blocks), 0.0f);
                if (tap < block_size) {
                    ggml_tensor * previous = ggml_view_3d(ctx0, blocks, hidden_size, block_size - tap, n_blocks,
                            blocks->nb[1], blocks->nb[2], 0);
                    values = ggml_concat(ctx0, zeros, previous, 1);
                } else {
                    values = zeros;
                }
            } else if (variant == V_PR && tap >= block_size) {
                values = ggml_fill(ctx0,
                        ggml_new_tensor_3d(ctx0, hidden->type, hidden_size, block_size, n_blocks), 0.0f);
            } else {
                ggml_tensor * previous = ggml_view_3d(ctx0, blocks, hidden_size, block_size - tap, n_blocks,
                        blocks->nb[1], blocks->nb[2], 0);
                values = ggml_pad_ext(ctx0, previous, 0, 0, tap, 0, 0, 0, 0, 0);
            }
        }
        values = ggml_reshape_2d(ctx0, values, hidden_size, n_tokens);
        ggml_tensor * weight = ggml_reshape_2d(ctx0,
                ggml_cont(ctx0, ggml_view_4d(ctx0, weight_all, group_size, n_groups, 1, n_tokens,
                        weight_all->nb[1], weight_all->nb[2], weight_all->nb[3], tap * weight_all->nb[2])),
                hidden_size, n_tokens);
        ggml_tensor * term = ggml_mul(ctx0, weight, values);
        result = result ? ggml_add(ctx0, result, term) : term;
    }
    return result;
}

static std::vector<float> run(ggml_backend_t be, const std::vector<float> & h, const std::vector<float> & w,
        int64_t H, int64_t T, int64_t G, int64_t K, int64_t NB, int variant, int * n_nodes) {
    ggml_init_params ip = { 64*1024*1024, nullptr, true };
    ggml_context * ctx = ggml_init(ip);
    ggml_tensor * hidden = ggml_new_tensor_2d(ctx, GGML_TYPE_F32, H, T);
    ggml_tensor * wall   = ggml_new_tensor_4d(ctx, GGML_TYPE_F32, G, H / G, K, T);
    ggml_tensor * out    = build(ctx, hidden, wall, K, NB, variant);
    ggml_cgraph * gf = ggml_new_graph(ctx);
    ggml_build_forward_expand(gf, out);
    *n_nodes = ggml_graph_n_nodes(gf);
    ggml_backend_buffer_t buf = ggml_backend_alloc_ctx_tensors(ctx, be);
    ggml_backend_tensor_set(hidden, h.data(), 0, h.size()*4);
    ggml_backend_tensor_set(wall,   w.data(), 0, w.size()*4);
    ggml_backend_graph_compute(be, gf);
    std::vector<float> r(ggml_nelements(out));
    ggml_backend_tensor_get(out, r.data(), 0, r.size()*4);
    ggml_backend_buffer_free(buf);
    ggml_free(ctx);
    return r;
}

int main() {
    ggml_backend_load_all();
    std::vector<ggml_backend_t> bes;
    for (size_t i = 0; i < ggml_backend_dev_count(); ++i) {
        ggml_backend_dev_t d = ggml_backend_dev_get(i);
        if (ggml_backend_dev_type(d) != GGML_BACKEND_DEVICE_TYPE_ACCEL) {
            bes.push_back(ggml_backend_dev_init(d, nullptr));
        }
    }
    // {hidden, group, kernel, block_size, n_blocks}
    const int64_t cases[][5] = {
        {64, 16, 4, 16, 3}, {64, 16, 4, 4, 2}, {64, 16, 4, 3, 5}, {64, 16, 4, 1, 7},
        {128, 32, 7, 2, 4}, {256, 64, 3, 16, 1}, {32, 8, 5, 5, 1},
    };
    std::mt19937 rng(42);
    std::normal_distribution<float> nd;
    int fails = 0;
    for (auto & c : cases) {
        const int64_t H = c[0], G = c[1], K = c[2], BS = c[3], NB = c[4], T = BS*NB;
        std::vector<float> h(H*T), w(G*(H/G)*K*T);
        for (auto & x : h) x = nd(rng);
        for (auto & x : w) x = nd(rng);
        for (auto be : bes) {
            int nm, np, nn;
            auto rm = run(be, h, w, H, T, G, K, NB, V_MASTER, &nm);
            auto rp = run(be, h, w, H, T, G, K, NB, V_PR,     &np);
            auto rn = run(be, h, w, H, T, G, K, NB, V_NEW,    &nn);
            const bool okp = memcmp(rm.data(), rp.data(), rm.size()*4) == 0;
            const bool okn = memcmp(rm.data(), rn.data(), rm.size()*4) == 0;
            fails += !okp + !okn;
            printf("%-6s H=%-3lld K=%lld block=%-2lld blocks=%lld  nodes master/pr/new %2d/%2d/%2d  pr %s  new %s\n",
                ggml_backend_name(be), (long long)H, (long long)K, (long long)BS, (long long)NB, nm, np, nn,
                okp ? "bit-exact" : "DIFF", okn ? "bit-exact" : "DIFF");
        }
    }
    printf("%s\n", fails ? "FAIL" : "ALL BIT-EXACT");
    for (auto be : bes) ggml_backend_free(be);
    return fails != 0;
}

@ServeurpersoCom
ServeurpersoCom merged commit 57b557c into ggml-org:master Sep 28, 2026
13 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
* models: pad on the left with ggml_pad_ext

The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders
build a left padding as a right pad followed by a roll, and DFlash2
concatenates a zero filled block in front of the previous tokens.
ggml_pad_ext does both in one node now that every backend supports a
left padding. The Gemma 4 audio embeddings are bit identical.

* models: skip the DFlash2 taps that only read padding

A tap at or past block_size shifts every row out of the block, so its
term is zero. The loop runs min(kernel_size, block_size) taps.
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* models: pad on the left with ggml_pad_ext

The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders
build a left padding as a right pad followed by a roll, and DFlash2
concatenates a zero filled block in front of the previous tokens.
ggml_pad_ext does both in one node now that every backend supports a
left padding. The Gemma 4 audio embeddings are bit identical.

* models: skip the DFlash2 taps that only read padding

A tap at or past block_size shifts every row out of the block, so its
term is zero. The loop runs min(kernel_size, block_size) taps.
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
* models: pad on the left with ggml_pad_ext

The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders
build a left padding as a right pad followed by a roll, and DFlash2
concatenates a zero filled block in front of the previous tokens.
ggml_pad_ext does both in one node now that every backend supports a
left padding. The Gemma 4 audio embeddings are bit identical.

* models: skip the DFlash2 taps that only read padding

A tap at or past block_size shifts every row out of the block, so its
term is zero. The loop runs min(kernel_size, block_size) taps.

(cherry picked from commit 57b557c)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model Model specific mtmd Related to multimodal functionality (video/image/audio)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants