Skip to content

ggml-cpu: AVX2 / AVX-VNNI GEMV+GEMM kernels for Q1_0 (4x8 repack) + shared-scale Q8 activations - #233

Merged
bri-prism merged 1 commit into
PrismML-Eng:prismfrom
Strike3D:q1_0-avx2-kernels
Sep 22, 2026
Merged

bri-prism merged 1 commit into
PrismML-Eng:prismfrom
Strike3D:q1_0-avx2-kernels

Conversation

@Strike3D

Copy link
Copy Markdown

ggml-cpu: AVX2/AVX-VNNI kernels for Q1_0 4x8 repack path

On x86 the fast Q1_0 (4x8 repack) path was only enabled with AVX-512 VNNI.
CPUs without AVX-512 (Intel Core Ultra / Alder..Arrow Lake, AMD Zen 2/3,
older Intel) fell back to the per-row vec_dot, which is compute-bound
(~11 ops per 32 weights) and leaves most of the memory bandwidth unused.

This adds:
- AVX2 GEMV and GEMM kernels for block_q1_0x4 (8-byte interleave), using
  AVX-VNNI vpdpbusd when available and pmaddubsw otherwise. Sign bits are
  kept in place as byte factors 2^q and multiplied directly; lanes are
  arranged so each dword sums equal factors, and the 2^q normalization is
  a single arithmetic shift per 128-weight block instead of per 32.
- Q8_0 activation quantizers with one scale per 128 values (aligned to
  QK1_0): ggml_quantize_row_q8_0_g128 / ggml_quantize_mat_q8_0_4x8_g128.
  Layout is unchanged (block_q8_0 / block_q8_0x4), so the existing
  AVX-512 and NEON kernels keep working; the new kernels accumulate a
  whole block in int32 and convert to float once. Precision matches Q8_K.
- Repack selection: Q1_0 now uses the 4x8 path on any AVX2 CPU.
- GGML_Q1_0_DISABLE_AVX512: test-only switch to exercise the AVX2 path on
  AVX-512 machines.

Measured (see PR description): kernel microbench 1.75-2.6x GEMV,
2.0-2.9x GEMM; end-to-end prompt processing 1.4-1.9x, generation
1.0-1.27x (memory-bandwidth bound after the change) on Intel Core
Ultra 7 255H and AMD Ryzen 9 3900X, Clang and MSVC.

Descrizione della PR

Problem

The Q1_0 (Bonsai) repack path (q1_0_4x8_q8_0) is selected on x86 only when
ggml_cpu_has_avx512() && ggml_cpu_has_avx512_vnni(). Every other x86 CPU
— including all recent Intel consumer parts (Alder Lake → Arrow Lake / Core
Ultra, which have no AVX-512) and AMD Zen 2/3 — goes through
ggml_vec_dot_q1_0_q8_0. That kernel is compute-bound (~11 instructions per
32 weights: bit expansion into byte masks, sign xor, sum, int→float, FMA),
so on these CPUs Bonsai-27B runs at ~5.5 tok/s while using roughly a quarter of
the available memory bandwidth.

Change

  1. AVX2 GEMV/GEMM for block_q1_0x4 (ggml/src/ggml-cpu/arch/x86/repack.cpp).
    AVX2 has no mask registers, so expanding 32 sign bits into 32 byte masks
    costs 4 ops. Instead, each bit is kept where it is as the byte factor
    2^q and multiplied directly with vpdpbusd (AVX-VNNI / AVX512-VL VNNI)
    or pmaddubsw+pmaddwd (plain AVX2). Lanes are arranged as
    L = 4*q + c so that the four bytes summed by one dword all carry the
    same factor; dividing by 2^q is one vpsravd per 128-weight block.
    With P = sum of activations over positive weights and S = total,
    dot = 2P - S. Activations are permuted once per call to match.
    GEMM shares the weight expansion across 4 activation rows.

  2. Shared-scale Q8 activations (repack.cpp, repack.h):
    ggml_quantize_row_q8_0_g128 and ggml_quantize_mat_q8_0_4x8_g128
    quantize with one scale per 128 values (= QK1_0). The block layout is
    the plain block_q8_0 / block_q8_0x4, so every existing kernel still
    works unchanged; the new kernels rely on the four scales of a group being
    equal and convert int→float once per block instead of once per 32. This
    is the precision of Q8_K (256-group scales), used throughout llama.cpp.
    Dispatched via ggml_repack_quantize_row/mat<BLOC_TYPE,...>, specialized
    for block_q1_0 only.

  3. Selection: Q1_0 → q1_0_4x8_q8_0 on any CPU with AVX2.

  4. GGML_Q1_0_DISABLE_AVX512: compile-time switch (off by default) to force
    the AVX2 path on AVX-512 machines, for testing/benchmarking only.

Correctness

Standalone harness (q1test.cpp, not included in the PR; happy to add it as a
tests/ target if wanted) compares the new kernels against the generic
*_generic kernels and against a double-precision reference computed from
the quantized integers, for n ∈ {128 … 16640}, 1–16 rows, normal and
heavy-tailed scales, both the VNNI and the pmaddubsw variants.
Max relative error: new kernels 2.1e-5, upstream generic kernels 3.4e-4.
Greedy (temperature 0) outputs of Bonsai-27B are unchanged in practice.

Performance

Kernel microbenchmark, single thread (Xeon, Skylake-SP; ratios vs the current
vec_dot path, n=5120, nc=5120):

pmaddubsw (AVX2) vpdpbusd (VNNI)
GEMV 1.75x 2.6x
GEMM (16 rows) 2.0x 2.9x

End-to-end, llama-bench -p 512 -n 128 on Bonsai-27B-Q1_0 (tok/s). Every row
is the same build with and without this patch.

Intel Core Ultra 7 255H (Arrow Lake-H, 6P+8E+2LPE, no AVX-512, AVX-VNNI;
32 GB DDR5-5600 single channel), Clang 19, -march=native:

threads pp512 before pp512 after tg128 before tg128 after
6 8.22 14.28 (1.74x) 5.34 6.47 (1.21x)
8 8.85 16.07 (1.82x) 5.50 6.81 (1.24x)
16 10.34 18.53 (1.79x) 5.19 6.38 (1.23x)

AMD Ryzen 9 3900X (Zen 2, 12C/24T, AVX2 without VNNI; 64 GB DDR4-3200
dual channel), three build recipes:

build threads pp512 before → after tg128 before → after
Clang 19, -march=native, no OpenMP 6 5.43 → 10.29 (1.90x) 4.37 → 5.55 (1.27x)
12 8.62 → 15.70 (1.82x) 6.00 → 6.72 (1.12x)
MSVC 19.44 + OpenMP, GGML_CPU_ALL_VARIANTS 6 6.74 → 9.83 (1.46x) 5.17 → 5.65 (1.09x)
12 10.54 → 14.33 (1.36x) 6.06 → 6.48 (1.07x)
CI recipe: Clang 19 + libomp, GGML_CPU_ALL_VARIANTS (haswell variant) 6 5.91 → 10.48 (1.77x) 4.66 → 5.52 (1.18x)
12 9.99 → 15.83 (1.58x) 6.50 → 6.63 (1.02x)

Observations:

  • Prompt processing gains 1.4–1.9x everywhere.
  • Generation gains 1.1–1.27x and then flattens: after the patch, tg is
    memory-bandwidth bound (thread scaling saturates at 3–4 threads on the
    laptop, and at 12 threads on the 3900X the gain vanishes). This is the
    expected regime for a 3.5 GiB weight stream; the kernel is no longer the
    bottleneck.
  • MSVC produces noticeably slower code for these kernels than Clang (1.4x vs
    1.9x on pp). Not a regression, but there is room to help MSVC's codegen.

Notes for reviewers

  • The g128 activation quantizer also feeds the existing AVX-512 and NEON
    Q1_0 kernels, which stay correct but see slightly coarser activation
    scales (one per 128 instead of per 32). If you prefer to keep those paths
    numerically untouched, the g128 quantizer can be limited to the AVX2 path
    by giving it a dedicated traits type; I chose the minimal change.
  • MSVC does not define __F16C__; the fp16 column-scale conversion falls
    back to scalar there (4 conversions per 128-weight block, negligible).
  • GEMM allocates a per-call scratch buffer (4*n bytes + scales) with
    malloc; GEMV uses a 16 KiB stack buffer with heap fallback.

Tested on Windows 11 with Clang 19.1.5 (-march=native and the CI
GGML_CPU_ALL_VARIANTS recipe) and MSVC 19.44, and on Linux with GCC 13
(AVX2 and AVX512-VL+VNNI builds).

## Titolo

    ggml-cpu: AVX2 / AVX-VNNI GEMV+GEMM kernels for Q1_0 (4x8 repack) + shared-scale Q8 activations

## Messaggio di commit (prima riga = titolo, poi corpo)

    ggml-cpu: AVX2/AVX-VNNI kernels for Q1_0 4x8 repack path

    On x86 the fast Q1_0 (4x8 repack) path was only enabled with AVX-512 VNNI.
    CPUs without AVX-512 (Intel Core Ultra / Alder..Arrow Lake, AMD Zen 2/3,
    older Intel) fell back to the per-row vec_dot, which is compute-bound
    (~11 ops per 32 weights) and leaves most of the memory bandwidth unused.

    This adds:
    - AVX2 GEMV and GEMM kernels for block_q1_0x4 (8-byte interleave), using
      AVX-VNNI vpdpbusd when available and pmaddubsw otherwise. Sign bits are
      kept in place as byte factors 2^q and multiplied directly; lanes are
      arranged so each dword sums equal factors, and the 2^q normalization is
      a single arithmetic shift per 128-weight block instead of per 32.
    - Q8_0 activation quantizers with one scale per 128 values (aligned to
      QK1_0): ggml_quantize_row_q8_0_g128 / ggml_quantize_mat_q8_0_4x8_g128.
      Layout is unchanged (block_q8_0 / block_q8_0x4), so the existing
      AVX-512 and NEON kernels keep working; the new kernels accumulate a
      whole block in int32 and convert to float once. Precision matches Q8_K.
    - Repack selection: Q1_0 now uses the 4x8 path on any AVX2 CPU.
    - GGML_Q1_0_DISABLE_AVX512: test-only switch to exercise the AVX2 path on
      AVX-512 machines.

    Measured (see PR description): kernel microbench 1.75-2.6x GEMV,
    2.0-2.9x GEMM; end-to-end prompt processing 1.4-1.9x, generation
    1.0-1.27x (memory-bandwidth bound after the change) on Intel Core
    Ultra 7 255H and AMD Ryzen 9 3900X, Clang and MSVC.

## Descrizione della PR

### Problem

The Q1_0 (Bonsai) repack path (`q1_0_4x8_q8_0`) is selected on x86 only when
`ggml_cpu_has_avx512() && ggml_cpu_has_avx512_vnni()`. Every other x86 CPU
— including all recent Intel consumer parts (Alder Lake → Arrow Lake / Core
Ultra, which have no AVX-512) and AMD Zen 2/3 — goes through
`ggml_vec_dot_q1_0_q8_0`. That kernel is compute-bound (~11 instructions per
32 weights: bit expansion into byte masks, sign xor, sum, int→float, FMA),
so on these CPUs Bonsai-27B runs at ~5.5 tok/s while using roughly a quarter of
the available memory bandwidth.

### Change

1. **AVX2 GEMV/GEMM for `block_q1_0x4`** (`ggml/src/ggml-cpu/arch/x86/repack.cpp`).
   AVX2 has no mask registers, so expanding 32 sign bits into 32 byte masks
   costs 4 ops. Instead, each bit is kept where it is as the byte factor
   `2^q` and multiplied directly with `vpdpbusd` (AVX-VNNI / AVX512-VL VNNI)
   or `pmaddubsw`+`pmaddwd` (plain AVX2). Lanes are arranged as
   `L = 4*q + c` so that the four bytes summed by one dword all carry the
   same factor; dividing by `2^q` is one `vpsravd` per 128-weight block.
   With `P` = sum of activations over positive weights and `S` = total,
   `dot = 2P - S`. Activations are permuted once per call to match.
   GEMM shares the weight expansion across 4 activation rows.

2. **Shared-scale Q8 activations** (`repack.cpp`, `repack.h`):
   `ggml_quantize_row_q8_0_g128` and `ggml_quantize_mat_q8_0_4x8_g128`
   quantize with one scale per 128 values (= `QK1_0`). The block layout is
   the plain `block_q8_0` / `block_q8_0x4`, so every existing kernel still
   works unchanged; the new kernels rely on the four scales of a group being
   equal and convert int→float once per block instead of once per 32. This
   is the precision of Q8_K (256-group scales), used throughout llama.cpp.
   Dispatched via `ggml_repack_quantize_row/mat<BLOC_TYPE,...>`, specialized
   for `block_q1_0` only.

3. **Selection**: `Q1_0` → `q1_0_4x8_q8_0` on any CPU with AVX2.

4. `GGML_Q1_0_DISABLE_AVX512`: compile-time switch (off by default) to force
   the AVX2 path on AVX-512 machines, for testing/benchmarking only.

### Correctness

Standalone harness (`q1test.cpp`, not included in the PR; happy to add it as a
`tests/` target if wanted) compares the new kernels against the generic
`*_generic` kernels and against a double-precision reference computed from
the quantized integers, for n ∈ {128 … 16640}, 1–16 rows, normal and
heavy-tailed scales, both the VNNI and the pmaddubsw variants.
Max relative error: new kernels 2.1e-5, upstream generic kernels 3.4e-4.
Greedy (temperature 0) outputs of Bonsai-27B are unchanged in practice.

### Performance

Kernel microbenchmark, single thread (Xeon, Skylake-SP; ratios vs the current
`vec_dot` path, n=5120, nc=5120):

| | pmaddubsw (AVX2) | vpdpbusd (VNNI) |
|---|---|---|
| GEMV | 1.75x | 2.6x |
| GEMM (16 rows) | 2.0x | 2.9x |

End-to-end, `llama-bench -p 512 -n 128` on Bonsai-27B-Q1_0 (tok/s). Every row
is the same build with and without this patch.

**Intel Core Ultra 7 255H** (Arrow Lake-H, 6P+8E+2LPE, no AVX-512, AVX-VNNI;
32 GB DDR5-5600 single channel), Clang 19, `-march=native`:

| threads | pp512 before | pp512 after | tg128 before | tg128 after |
|---|---|---|---|---|
| 6 | 8.22 | 14.28 (1.74x) | 5.34 | 6.47 (1.21x) |
| 8 | 8.85 | 16.07 (1.82x) | 5.50 | 6.81 (1.24x) |
| 16 | 10.34 | 18.53 (1.79x) | 5.19 | 6.38 (1.23x) |

**AMD Ryzen 9 3900X** (Zen 2, 12C/24T, AVX2 without VNNI; 64 GB DDR4-3200
dual channel), three build recipes:

| build | threads | pp512 before → after | tg128 before → after |
|---|---|---|---|
| Clang 19, `-march=native`, no OpenMP | 6 | 5.43 → 10.29 (1.90x) | 4.37 → 5.55 (1.27x) |
| | 12 | 8.62 → 15.70 (1.82x) | 6.00 → 6.72 (1.12x) |
| MSVC 19.44 + OpenMP, `GGML_CPU_ALL_VARIANTS` | 6 | 6.74 → 9.83 (1.46x) | 5.17 → 5.65 (1.09x) |
| | 12 | 10.54 → 14.33 (1.36x) | 6.06 → 6.48 (1.07x) |
| **CI recipe**: Clang 19 + libomp, `GGML_CPU_ALL_VARIANTS` (haswell variant) | 6 | 5.91 → 10.48 (1.77x) | 4.66 → 5.52 (1.18x) |
| | 12 | 9.99 → 15.83 (1.58x) | 6.50 → 6.63 (1.02x) |

Observations:
- Prompt processing gains 1.4–1.9x everywhere.
- Generation gains 1.1–1.27x and then flattens: after the patch, tg is
  memory-bandwidth bound (thread scaling saturates at 3–4 threads on the
  laptop, and at 12 threads on the 3900X the gain vanishes). This is the
  expected regime for a 3.5 GiB weight stream; the kernel is no longer the
  bottleneck.
- MSVC produces noticeably slower code for these kernels than Clang (1.4x vs
  1.9x on pp). Not a regression, but there is room to help MSVC's codegen.

### Notes for reviewers

- The g128 activation quantizer also feeds the existing AVX-512 and NEON
  Q1_0 kernels, which stay correct but see slightly coarser activation
  scales (one per 128 instead of per 32). If you prefer to keep those paths
  numerically untouched, the g128 quantizer can be limited to the AVX2 path
  by giving it a dedicated traits type; I chose the minimal change.
- MSVC does not define `__F16C__`; the fp16 column-scale conversion falls
  back to scalar there (4 conversions per 128-weight block, negligible).
- GEMM allocates a per-call scratch buffer (`4*n` bytes + scales) with
  `malloc`; GEMV uses a 16 KiB stack buffer with heap fallback.

Tested on Windows 11 with Clang 19.1.5 (`-march=native` and the CI
`GGML_CPU_ALL_VARIANTS` recipe) and MSVC 19.44, and on Linux with GCC 13
(AVX2 and AVX512-VL+VNNI builds).
@github-actions github-actions Bot added the ggml label Sep 21, 2026
@bri-prism

Copy link
Copy Markdown
Collaborator

Confirmed on the exact CPU class this targets — Intel Core Ultra X7 358H (AVX2 + AVX-VNNI, no AVX-512)

This machine is the case in your description: AVX2=1, AVX512F=0, AVX_VNNI=1, 16 cores, Windows 11. So it takes the new vpdpbusd path rather than the pmaddubsw fallback, and before this PR it was on the per-row vec_dot.

Builds: prism @ 422590f5d vs this PR @ 538bb3fdf, MinGW/ucrt64 g++, Ninja, Release, -DGGML_NATIVE=ON -DGGML_VULKAN=OFF. Model Bonsai-1.7B-Q1_0.gguf, everything on CPU (-ngl 0), machine otherwise idle.

Throughput — ABAB interleaved, two rounds

llama-bench -p 512 -n 128 -ngl 0 -r 3

round arm pp512 tg128
1 prism 151.11 ± 1.72 24.01 ± 0.13
1 #233 302.59 ± 2.14 26.51 ± 0.22
2 prism 135.97 ± 8.13 23.94 ± 0.19
2 #233 266.70 ± 8.86 27.15 ± 0.03

Means: pp512 143.5 -> 284.6 = 1.98x, tg128 23.98 -> 26.83 = 1.12x. Prefill sits at the top of your stated 1.4-1.9x band; generation is mid-range of 1.0-1.27x, consistent with it being bandwidth-bound.

Quality — the shared-scale Q8 activations do not cost anything measurable

This is the part I most wanted to check, since the per-128 shared scale is a real numerical change rather than a pure speed one. llama-perplexity -f wiki.test.raw -c 512 --chunks 40, same model, same machine:

build PPL
prism 422590f5d 22.0195 +/- 0.76303
#233 538bb3fdf 22.0078 +/- 0.76245

Statistically indistinguishable — the PR is nominally 0.05% lower, three orders of magnitude inside the error bar, and the per-chunk series track each other to four significant figures at every one of the 40 chunks. Your "precision matches Q8_K" claim holds on this model.

Greedy generation (--temp 0 --seed 1234, 80 tokens, same prompt) diverges from the baseline at exactly one clause — "Instead of decoding one token at a time" vs "Instead of decoding the entire sequence at once" — with identical text either side of it. That is a single near-tie token flip, which is what a numerics change should look like; both readings are coherent.

Two things worth fixing before merge

1. test-backend-ops does not cover this PR, and it looks like it does. test-backend-ops test -o MUL_MAT -b CPU reports 1294/1294 passed on this branch — but the repack kernels are reached through ggml-cpu's extra buffer type during model load, which that suite does not route through. A reviewer could easily read that green run as validation when it exercises none of the new code. The evidence that actually covers this PR is the perplexity and generation comparison above. Might be worth saying so in the description, or adding a test that allocates through the repack buffer type.

2. The commit message is the PR-drafting notes. The single commit's message begins:

# Pull request — testo pronto da incollare su GitHub

## Titolo
    ggml-cpu: AVX2 / AVX-VNNI GEMV+GEMM kernels for Q1_0 ...
## Messaggio di commit (prima riga = titolo, poi corpo)
    ggml-cpu: AVX2/AVX-VNNI kernels for Q1_0 4x8 repack path
    ...

The intended message is nested inside the scaffolding. Worth a rewrite so the history reads properly.

Scope of what I ran

1.7B only — I picked it so perplexity was affordable on CPU. I have Bonsai-27B-Q1_0 here too and can run the same pair at 27B if a larger-model data point would help; say the word.

@bri-prism

Copy link
Copy Markdown
Collaborator

27B data point, as offered — and the gains are larger than at 1.7B, above your stated ranges

Same machine (Core Ultra X7 358H, AVX2 + AVX-VNNI, no AVX-512), same builds (prism 422590f5d vs this PR 538bb3fdf), Bonsai-27B-Q1_0.gguf (3.54 GiB), everything on CPU, machine idle and runs strictly serialized.

Throughput — ABAB, two rounds

round arm pp512 tg128
1 prism 7.20 ± 0.30 3.51 ± 0.02
1 #233 16.60 ± 0.16 4.88 ± 0.01
2 prism 7.07 ± 0.01 3.50 ± 0.00
2 #233 16.78 ± 0.01 4.89 ± 0.04

Means: pp512 7.14 -> 16.69 = 2.34x, tg128 3.505 -> 4.885 = 1.39x.

Side by side with the 1.7B numbers I posted earlier, same machine and method:

model pp512 tg128
Bonsai-1.7B-Q1_0 143.5 -> 284.6 = 1.98x 23.98 -> 26.83 = 1.12x
Bonsai-27B-Q1_0 7.14 -> 16.69 = 2.34x 3.505 -> 4.885 = 1.39x

Both gains grow with model size, and both 27B figures sit above the ranges in your description (1.4-1.9x prefill, 1.0-1.27x generation). The generation one is the notable miss: 1.39x at 27B against a stated ceiling of 1.27x.

Worth saying that I predicted the opposite before running it — I expected the generation gain to shrink at 27B, reasoning that 3.54 GiB falls out of cache into a bandwidth-bound regime where a compute-side kernel win stops mattering. That was wrong, and consistently so across both rounds (spread ±0.01-0.04 on the tg rows). Whatever the mechanism, the per-row vec_dot baseline appears to be costing more at 27B than a purely bandwidth-bound model predicts. You may want to re-measure on your own hardware before widening the claimed range, but on this CPU the PR is underselling itself.

Quality

llama-perplexity -f wiki.test.raw -c 512 --chunks 40, same model and machine:

build PPL
prism 422590f5d 11.5899 +/- 0.32558
#233 538bb3fdf 11.5971 +/- 0.32586

+0.062%, two orders of magnitude inside the error bar.

Taken together with the 1.7B pair, where this PR came out 0.053% lower: the sign of the difference flips between models, which is what a numerically-neutral change should look like — tiny deviations in both directions rather than a consistent drift. I would read the two pairs together as better evidence for "precision matches Q8_K" than either one alone.

Both caveats from my earlier comment still stand: test-backend-ops does not exercise the repack buffer type and so is not coverage for this PR, and the commit message is still the PR-drafting scaffolding.

@bri-prism
bri-prism self-requested a review September 22, 2026 04:02
@bri-prism
bri-prism merged commit ea50aba into PrismML-Eng:prism Sep 22, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants