Skip to content

ggml-cpu: no repack/GEMM path for PTQ1_0 on x86, CPU prompt processing runs at decode speed #324

Description

@cheese-cakee

Prerequisites

Feature Description

On CPU, PTQ1_0 has only a row-by-row dot: nrows = 1, Q8_0 activations, and no entry in repack.cpp / arch/x86/repack.cpp. Prompt processing therefore runs one batch-one dot per token, and each packed block is decoded again for every token. Prefill runs at about decode speed.

i5-13450HX (AVX2 + AVX-VNNI, no AVX-512), WSL2, CPU-only Release build with GGML_NATIVE, 8 threads, llama-bench -p 512 -n 0 and -p 0 -n 32, -r 1 (single runs: direction only, not a benchmark claim). Profile: perf record -e cpu-clock (WSL2 has no hardware counters).

model / build pp512 t/s tg32 t/s top symbol in pp512
Ternary-Bonsai-2-27B PTQ1_0, prism (SSE dot from #248) 1.95 1.59 ggml_vec_dot_ptq1_0_q8_0 98.1%
same, + #250 AVX2/VNNI dot (local rebase) 3.92 3.21 ggml_vec_dot_ptq1_0_q8_0 96.6%
same weights requantized to PQ2_0, prism 5.32 5.15 tinyBLAS_PQ2K_AVX::gemm<2,4> 95.0%
Qwen3-0.6B Q8_0, prism (-r 3, reference) 373 69.7 -

Rough throughput (2 x matmul params x pp t/s): PTQ1_0 + #250 about 200 GOPS, Q8_0 about 560 GOPS. So about 2.8x headroom on this CPU, with the caveat that the 0.6B model has a smaller working set.

test-quantize-perf --op vec_dot_q, single thread, cycles per 32 values: ptq1_0 (#250) ~3.95, q8_0 ~2.6-3.1. Trit unpack is about 35-40% of the PTQ1_0 dot. Removing only the repeated unpack caps the gain near 1.5x; the rest needs register tiling with weight and activation reuse, as in the Q8_0/Q4_0 GEMMs.

Motivation

Possible Implementation

x86 AVX2/AVX-VNNI first, in the pattern of #233: repack into interleaved blocks at load, GEMV + GEMM in arch/x86/repack.cpp, current vec_dot stays the fallback for other ISAs and shapes. #250's dot is the batch-one baseline, so this would come after #250 lands. Tests through existing test-backend-ops and test-quantize-fns, no new files under tests/.

Questions before I write code:

  1. Load-time repack (load time and memory cost on a 27B model) or an on-the-fly 2-4 column tile with no repack?
  2. Q8_0 or Q8_K activations?
  3. The PQ2_0 gemm<2,4> only reaches pp ~ tg on this CPU, so I would target the Q8_0 GEMM numbers, not PQ2_0. Does that match your view?

AI was used to help write the profiling scripts and to summarize the measurements; I ran and checked the numbers myself.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions