Prerequisites
Feature Description
On CPU, PTQ1_0 has only a row-by-row dot: nrows = 1, Q8_0 activations, and no entry in repack.cpp / arch/x86/repack.cpp. Prompt processing therefore runs one batch-one dot per token, and each packed block is decoded again for every token. Prefill runs at about decode speed.
i5-13450HX (AVX2 + AVX-VNNI, no AVX-512), WSL2, CPU-only Release build with GGML_NATIVE, 8 threads, llama-bench -p 512 -n 0 and -p 0 -n 32, -r 1 (single runs: direction only, not a benchmark claim). Profile: perf record -e cpu-clock (WSL2 has no hardware counters).
| model / build |
pp512 t/s |
tg32 t/s |
top symbol in pp512 |
Ternary-Bonsai-2-27B PTQ1_0, prism (SSE dot from #248) |
1.95 |
1.59 |
ggml_vec_dot_ptq1_0_q8_0 98.1% |
| same, + #250 AVX2/VNNI dot (local rebase) |
3.92 |
3.21 |
ggml_vec_dot_ptq1_0_q8_0 96.6% |
same weights requantized to PQ2_0, prism |
5.32 |
5.15 |
tinyBLAS_PQ2K_AVX::gemm<2,4> 95.0% |
Qwen3-0.6B Q8_0, prism (-r 3, reference) |
373 |
69.7 |
- |
Rough throughput (2 x matmul params x pp t/s): PTQ1_0 + #250 about 200 GOPS, Q8_0 about 560 GOPS. So about 2.8x headroom on this CPU, with the caveat that the 0.6B model has a smaller working set.
test-quantize-perf --op vec_dot_q, single thread, cycles per 32 values: ptq1_0 (#250) ~3.95, q8_0 ~2.6-3.1. Trit unpack is about 35-40% of the PTQ1_0 dot. Removing only the repeated unpack caps the gain near 1.5x; the rest needs register tiling with weight and activation reuse, as in the Q8_0/Q4_0 GEMMs.
Motivation
Possible Implementation
x86 AVX2/AVX-VNNI first, in the pattern of #233: repack into interleaved blocks at load, GEMV + GEMM in arch/x86/repack.cpp, current vec_dot stays the fallback for other ISAs and shapes. #250's dot is the batch-one baseline, so this would come after #250 lands. Tests through existing test-backend-ops and test-quantize-fns, no new files under tests/.
Questions before I write code:
- Load-time repack (load time and memory cost on a 27B model) or an on-the-fly 2-4 column tile with no repack?
- Q8_0 or Q8_K activations?
- The PQ2_0
gemm<2,4> only reaches pp ~ tg on this CPU, so I would target the Q8_0 GEMM numbers, not PQ2_0. Does that match your view?
AI was used to help write the profiling scripts and to summarize the measurements; I ran and checked the numbers myself.
Prerequisites
prism@ 6bfcd79. Numbers below are from 2459f68; no x86ggml-cpufile changed between the two.Feature Description
On CPU, PTQ1_0 has only a row-by-row dot:
nrows = 1, Q8_0 activations, and no entry inrepack.cpp/arch/x86/repack.cpp. Prompt processing therefore runs one batch-one dot per token, and each packed block is decoded again for every token. Prefill runs at about decode speed.i5-13450HX (AVX2 + AVX-VNNI, no AVX-512), WSL2, CPU-only Release build with
GGML_NATIVE, 8 threads,llama-bench -p 512 -n 0and-p 0 -n 32,-r 1(single runs: direction only, not a benchmark claim). Profile:perf record -e cpu-clock(WSL2 has no hardware counters).prism(SSE dot from #248)ggml_vec_dot_ptq1_0_q8_098.1%ggml_vec_dot_ptq1_0_q8_096.6%prismtinyBLAS_PQ2K_AVX::gemm<2,4>95.0%prism(-r 3, reference)Rough throughput (2 x matmul params x pp t/s): PTQ1_0 + #250 about 200 GOPS, Q8_0 about 560 GOPS. So about 2.8x headroom on this CPU, with the caveat that the 0.6B model has a smaller working set.
test-quantize-perf --op vec_dot_q, single thread, cycles per 32 values: ptq1_0 (#250) ~3.95, q8_0 ~2.6-3.1. Trit unpack is about 35-40% of the PTQ1_0 dot. Removing only the repeated unpack caps the gain near 1.5x; the rest needs register tiling with weight and activation reuse, as in the Q8_0/Q4_0 GEMMs.Motivation
Possible Implementation
x86 AVX2/AVX-VNNI first, in the pattern of #233: repack into interleaved blocks at load, GEMV + GEMM in
arch/x86/repack.cpp, current vec_dot stays the fallback for other ISAs and shapes. #250's dot is the batch-one baseline, so this would come after #250 lands. Tests through existingtest-backend-opsandtest-quantize-fns, no new files undertests/.Questions before I write code:
gemm<2,4>only reaches pp ~ tg on this CPU, so I would target the Q8_0 GEMM numbers, not PQ2_0. Does that match your view?AI was used to help write the profiling scripts and to summarize the measurements; I ran and checked the numbers myself.