Repository navigation
ggml-cpu: AVX2 / AVX-VNNI GEMV+GEMM kernels for Q1_0 (4x8 repack) + shared-scale Q8 activations - #233
Conversation
## Titolo
ggml-cpu: AVX2 / AVX-VNNI GEMV+GEMM kernels for Q1_0 (4x8 repack) + shared-scale Q8 activations
## Messaggio di commit (prima riga = titolo, poi corpo)
ggml-cpu: AVX2/AVX-VNNI kernels for Q1_0 4x8 repack path
On x86 the fast Q1_0 (4x8 repack) path was only enabled with AVX-512 VNNI.
CPUs without AVX-512 (Intel Core Ultra / Alder..Arrow Lake, AMD Zen 2/3,
older Intel) fell back to the per-row vec_dot, which is compute-bound
(~11 ops per 32 weights) and leaves most of the memory bandwidth unused.
This adds:
- AVX2 GEMV and GEMM kernels for block_q1_0x4 (8-byte interleave), using
AVX-VNNI vpdpbusd when available and pmaddubsw otherwise. Sign bits are
kept in place as byte factors 2^q and multiplied directly; lanes are
arranged so each dword sums equal factors, and the 2^q normalization is
a single arithmetic shift per 128-weight block instead of per 32.
- Q8_0 activation quantizers with one scale per 128 values (aligned to
QK1_0): ggml_quantize_row_q8_0_g128 / ggml_quantize_mat_q8_0_4x8_g128.
Layout is unchanged (block_q8_0 / block_q8_0x4), so the existing
AVX-512 and NEON kernels keep working; the new kernels accumulate a
whole block in int32 and convert to float once. Precision matches Q8_K.
- Repack selection: Q1_0 now uses the 4x8 path on any AVX2 CPU.
- GGML_Q1_0_DISABLE_AVX512: test-only switch to exercise the AVX2 path on
AVX-512 machines.
Measured (see PR description): kernel microbench 1.75-2.6x GEMV,
2.0-2.9x GEMM; end-to-end prompt processing 1.4-1.9x, generation
1.0-1.27x (memory-bandwidth bound after the change) on Intel Core
Ultra 7 255H and AMD Ryzen 9 3900X, Clang and MSVC.
## Descrizione della PR
### Problem
The Q1_0 (Bonsai) repack path (`q1_0_4x8_q8_0`) is selected on x86 only when
`ggml_cpu_has_avx512() && ggml_cpu_has_avx512_vnni()`. Every other x86 CPU
— including all recent Intel consumer parts (Alder Lake → Arrow Lake / Core
Ultra, which have no AVX-512) and AMD Zen 2/3 — goes through
`ggml_vec_dot_q1_0_q8_0`. That kernel is compute-bound (~11 instructions per
32 weights: bit expansion into byte masks, sign xor, sum, int→float, FMA),
so on these CPUs Bonsai-27B runs at ~5.5 tok/s while using roughly a quarter of
the available memory bandwidth.
### Change
1. **AVX2 GEMV/GEMM for `block_q1_0x4`** (`ggml/src/ggml-cpu/arch/x86/repack.cpp`).
AVX2 has no mask registers, so expanding 32 sign bits into 32 byte masks
costs 4 ops. Instead, each bit is kept where it is as the byte factor
`2^q` and multiplied directly with `vpdpbusd` (AVX-VNNI / AVX512-VL VNNI)
or `pmaddubsw`+`pmaddwd` (plain AVX2). Lanes are arranged as
`L = 4*q + c` so that the four bytes summed by one dword all carry the
same factor; dividing by `2^q` is one `vpsravd` per 128-weight block.
With `P` = sum of activations over positive weights and `S` = total,
`dot = 2P - S`. Activations are permuted once per call to match.
GEMM shares the weight expansion across 4 activation rows.
2. **Shared-scale Q8 activations** (`repack.cpp`, `repack.h`):
`ggml_quantize_row_q8_0_g128` and `ggml_quantize_mat_q8_0_4x8_g128`
quantize with one scale per 128 values (= `QK1_0`). The block layout is
the plain `block_q8_0` / `block_q8_0x4`, so every existing kernel still
works unchanged; the new kernels rely on the four scales of a group being
equal and convert int→float once per block instead of once per 32. This
is the precision of Q8_K (256-group scales), used throughout llama.cpp.
Dispatched via `ggml_repack_quantize_row/mat<BLOC_TYPE,...>`, specialized
for `block_q1_0` only.
3. **Selection**: `Q1_0` → `q1_0_4x8_q8_0` on any CPU with AVX2.
4. `GGML_Q1_0_DISABLE_AVX512`: compile-time switch (off by default) to force
the AVX2 path on AVX-512 machines, for testing/benchmarking only.
### Correctness
Standalone harness (`q1test.cpp`, not included in the PR; happy to add it as a
`tests/` target if wanted) compares the new kernels against the generic
`*_generic` kernels and against a double-precision reference computed from
the quantized integers, for n ∈ {128 … 16640}, 1–16 rows, normal and
heavy-tailed scales, both the VNNI and the pmaddubsw variants.
Max relative error: new kernels 2.1e-5, upstream generic kernels 3.4e-4.
Greedy (temperature 0) outputs of Bonsai-27B are unchanged in practice.
### Performance
Kernel microbenchmark, single thread (Xeon, Skylake-SP; ratios vs the current
`vec_dot` path, n=5120, nc=5120):
| | pmaddubsw (AVX2) | vpdpbusd (VNNI) |
|---|---|---|
| GEMV | 1.75x | 2.6x |
| GEMM (16 rows) | 2.0x | 2.9x |
End-to-end, `llama-bench -p 512 -n 128` on Bonsai-27B-Q1_0 (tok/s). Every row
is the same build with and without this patch.
**Intel Core Ultra 7 255H** (Arrow Lake-H, 6P+8E+2LPE, no AVX-512, AVX-VNNI;
32 GB DDR5-5600 single channel), Clang 19, `-march=native`:
| threads | pp512 before | pp512 after | tg128 before | tg128 after |
|---|---|---|---|---|
| 6 | 8.22 | 14.28 (1.74x) | 5.34 | 6.47 (1.21x) |
| 8 | 8.85 | 16.07 (1.82x) | 5.50 | 6.81 (1.24x) |
| 16 | 10.34 | 18.53 (1.79x) | 5.19 | 6.38 (1.23x) |
**AMD Ryzen 9 3900X** (Zen 2, 12C/24T, AVX2 without VNNI; 64 GB DDR4-3200
dual channel), three build recipes:
| build | threads | pp512 before → after | tg128 before → after |
|---|---|---|---|
| Clang 19, `-march=native`, no OpenMP | 6 | 5.43 → 10.29 (1.90x) | 4.37 → 5.55 (1.27x) |
| | 12 | 8.62 → 15.70 (1.82x) | 6.00 → 6.72 (1.12x) |
| MSVC 19.44 + OpenMP, `GGML_CPU_ALL_VARIANTS` | 6 | 6.74 → 9.83 (1.46x) | 5.17 → 5.65 (1.09x) |
| | 12 | 10.54 → 14.33 (1.36x) | 6.06 → 6.48 (1.07x) |
| **CI recipe**: Clang 19 + libomp, `GGML_CPU_ALL_VARIANTS` (haswell variant) | 6 | 5.91 → 10.48 (1.77x) | 4.66 → 5.52 (1.18x) |
| | 12 | 9.99 → 15.83 (1.58x) | 6.50 → 6.63 (1.02x) |
Observations:
- Prompt processing gains 1.4–1.9x everywhere.
- Generation gains 1.1–1.27x and then flattens: after the patch, tg is
memory-bandwidth bound (thread scaling saturates at 3–4 threads on the
laptop, and at 12 threads on the 3900X the gain vanishes). This is the
expected regime for a 3.5 GiB weight stream; the kernel is no longer the
bottleneck.
- MSVC produces noticeably slower code for these kernels than Clang (1.4x vs
1.9x on pp). Not a regression, but there is room to help MSVC's codegen.
### Notes for reviewers
- The g128 activation quantizer also feeds the existing AVX-512 and NEON
Q1_0 kernels, which stay correct but see slightly coarser activation
scales (one per 128 instead of per 32). If you prefer to keep those paths
numerically untouched, the g128 quantizer can be limited to the AVX2 path
by giving it a dedicated traits type; I chose the minimal change.
- MSVC does not define `__F16C__`; the fp16 column-scale conversion falls
back to scalar there (4 conversions per 128-weight block, negligible).
- GEMM allocates a per-call scratch buffer (`4*n` bytes + scales) with
`malloc`; GEMV uses a 16 KiB stack buffer with heap fallback.
Tested on Windows 11 with Clang 19.1.5 (`-march=native` and the CI
`GGML_CPU_ALL_VARIANTS` recipe) and MSVC 19.44, and on Linux with GCC 13
(AVX2 and AVX512-VL+VNNI builds).
Confirmed on the exact CPU class this targets — Intel Core Ultra X7 358H (AVX2 + AVX-VNNI, no AVX-512)This machine is the case in your description: Builds: Throughput — ABAB interleaved, two rounds
Means: pp512 143.5 -> 284.6 = 1.98x, tg128 23.98 -> 26.83 = 1.12x. Prefill sits at the top of your stated 1.4-1.9x band; generation is mid-range of 1.0-1.27x, consistent with it being bandwidth-bound. Quality — the shared-scale Q8 activations do not cost anything measurableThis is the part I most wanted to check, since the per-128 shared scale is a real numerical change rather than a pure speed one.
Statistically indistinguishable — the PR is nominally 0.05% lower, three orders of magnitude inside the error bar, and the per-chunk series track each other to four significant figures at every one of the 40 chunks. Your "precision matches Q8_K" claim holds on this model. Greedy generation ( Two things worth fixing before merge1. 2. The commit message is the PR-drafting notes. The single commit's message begins: The intended message is nested inside the scaffolding. Worth a rewrite so the history reads properly. Scope of what I ran1.7B only — I picked it so perplexity was affordable on CPU. I have |
27B data point, as offered — and the gains are larger than at 1.7B, above your stated rangesSame machine (Core Ultra X7 358H, AVX2 + AVX-VNNI, no AVX-512), same builds ( Throughput — ABAB, two rounds
Means: pp512 7.14 -> 16.69 = 2.34x, tg128 3.505 -> 4.885 = 1.39x. Side by side with the 1.7B numbers I posted earlier, same machine and method:
Both gains grow with model size, and both 27B figures sit above the ranges in your description (1.4-1.9x prefill, 1.0-1.27x generation). The generation one is the notable miss: 1.39x at 27B against a stated ceiling of 1.27x. Worth saying that I predicted the opposite before running it — I expected the generation gain to shrink at 27B, reasoning that 3.54 GiB falls out of cache into a bandwidth-bound regime where a compute-side kernel win stops mattering. That was wrong, and consistently so across both rounds (spread ±0.01-0.04 on the tg rows). Whatever the mechanism, the per-row Quality
+0.062%, two orders of magnitude inside the error bar. Taken together with the 1.7B pair, where this PR came out 0.053% lower: the sign of the difference flips between models, which is what a numerically-neutral change should look like — tiny deviations in both directions rather than a consistent drift. I would read the two pairs together as better evidence for "precision matches Q8_K" than either one alone. Both caveats from my earlier comment still stand: |
ggml-cpu: AVX2/AVX-VNNI kernels for Q1_0 4x8 repack path
Descrizione della PR
Problem
The Q1_0 (Bonsai) repack path (
q1_0_4x8_q8_0) is selected on x86 only whenggml_cpu_has_avx512() && ggml_cpu_has_avx512_vnni(). Every other x86 CPU— including all recent Intel consumer parts (Alder Lake → Arrow Lake / Core
Ultra, which have no AVX-512) and AMD Zen 2/3 — goes through
ggml_vec_dot_q1_0_q8_0. That kernel is compute-bound (~11 instructions per32 weights: bit expansion into byte masks, sign xor, sum, int→float, FMA),
so on these CPUs Bonsai-27B runs at ~5.5 tok/s while using roughly a quarter of
the available memory bandwidth.
Change
AVX2 GEMV/GEMM for
block_q1_0x4(ggml/src/ggml-cpu/arch/x86/repack.cpp).AVX2 has no mask registers, so expanding 32 sign bits into 32 byte masks
costs 4 ops. Instead, each bit is kept where it is as the byte factor
2^qand multiplied directly withvpdpbusd(AVX-VNNI / AVX512-VL VNNI)or
pmaddubsw+pmaddwd(plain AVX2). Lanes are arranged asL = 4*q + cso that the four bytes summed by one dword all carry thesame factor; dividing by
2^qis onevpsravdper 128-weight block.With
P= sum of activations over positive weights andS= total,dot = 2P - S. Activations are permuted once per call to match.GEMM shares the weight expansion across 4 activation rows.
Shared-scale Q8 activations (
repack.cpp,repack.h):ggml_quantize_row_q8_0_g128andggml_quantize_mat_q8_0_4x8_g128quantize with one scale per 128 values (=
QK1_0). The block layout isthe plain
block_q8_0/block_q8_0x4, so every existing kernel stillworks unchanged; the new kernels rely on the four scales of a group being
equal and convert int→float once per block instead of once per 32. This
is the precision of Q8_K (256-group scales), used throughout llama.cpp.
Dispatched via
ggml_repack_quantize_row/mat<BLOC_TYPE,...>, specializedfor
block_q1_0only.Selection:
Q1_0→q1_0_4x8_q8_0on any CPU with AVX2.GGML_Q1_0_DISABLE_AVX512: compile-time switch (off by default) to forcethe AVX2 path on AVX-512 machines, for testing/benchmarking only.
Correctness
Standalone harness (
q1test.cpp, not included in the PR; happy to add it as atests/target if wanted) compares the new kernels against the generic*_generickernels and against a double-precision reference computed fromthe quantized integers, for n ∈ {128 … 16640}, 1–16 rows, normal and
heavy-tailed scales, both the VNNI and the pmaddubsw variants.
Max relative error: new kernels 2.1e-5, upstream generic kernels 3.4e-4.
Greedy (temperature 0) outputs of Bonsai-27B are unchanged in practice.
Performance
Kernel microbenchmark, single thread (Xeon, Skylake-SP; ratios vs the current
vec_dotpath, n=5120, nc=5120):End-to-end,
llama-bench -p 512 -n 128on Bonsai-27B-Q1_0 (tok/s). Every rowis the same build with and without this patch.
Intel Core Ultra 7 255H (Arrow Lake-H, 6P+8E+2LPE, no AVX-512, AVX-VNNI;
32 GB DDR5-5600 single channel), Clang 19,
-march=native:AMD Ryzen 9 3900X (Zen 2, 12C/24T, AVX2 without VNNI; 64 GB DDR4-3200
dual channel), three build recipes:
-march=native, no OpenMPGGML_CPU_ALL_VARIANTSGGML_CPU_ALL_VARIANTS(haswell variant)Observations:
memory-bandwidth bound (thread scaling saturates at 3–4 threads on the
laptop, and at 12 threads on the 3900X the gain vanishes). This is the
expected regime for a 3.5 GiB weight stream; the kernel is no longer the
bottleneck.
1.9x on pp). Not a regression, but there is room to help MSVC's codegen.
Notes for reviewers
Q1_0 kernels, which stay correct but see slightly coarser activation
scales (one per 128 instead of per 32). If you prefer to keep those paths
numerically untouched, the g128 quantizer can be limited to the AVX2 path
by giving it a dedicated traits type; I chose the minimal change.
__F16C__; the fp16 column-scale conversion fallsback to scalar there (4 conversions per 128-weight block, negligible).
4*nbytes + scales) withmalloc; GEMV uses a 16 KiB stack buffer with heap fallback.Tested on Windows 11 with Clang 19.1.5 (
-march=nativeand the CIGGML_CPU_ALL_VARIANTSrecipe) and MSVC 19.44, and on Linux with GCC 13(AVX2 and AVX512-VL+VNNI builds).