metal : add PTQ1_0 mat-vec for two to four columns - #262
Conversation
Assisted-by: OpenAI Codex
bri-prism
left a comment
There was a problem hiding this comment.
Tested on an M5 Pro (Metal), merged onto current prism (bda59ea); the merge is clean. test-backend-ops MUL_MAT PTQ1_0 passes 172/172 with GGML_METAL_PTQ1_MULTICOL off and on. Output with the switch on matches off on Ternary Bonsai 2 27B PTQ1_0 (llama-perplexity, ubatch 4: mean KLD ~0, max 3.7e-5, same top token 100%).
llama-bench, 27B PTQ1_0, -fa 1, two rounds, t/s: pp2 10.6 -> 37.6, pp3 15.0 -> 34.0, pp4 19.3 -> 39.9; pp1 and tg32 unchanged (25.1 / 25.8), and with the switch off it matches the prism base exactly. Without this, a two-column verify on Metal runs at under half the single-token speed, which is why speculative decoding (DFlash2 #261, MTP) loses on Metal today. Given there's no regression on the off path, it may be worth making this the default once it's been checked on one more Apple generation. LGTM.
Overview
Add an opt-in Metal PTQ1_0/F32 mat-vec for 2-4 columns, sharing decoded weights across columns and activations across four output rows. Enable with
GGML_METAL_PTQ1_MULTICOL=1; default remains off. Existing single-vector and larger-batch paths are unchanged. Five files; no benchmark framework, MTP tooling or device-specific dispatch.Single-user Bonsai 2 PTQ1_0 native generation tok/s:
For an M5 single user, candidate MTP is about 32% faster generation than candidate plain. M5 full-request throughput, including prompt/request overhead, is 27.96 plain and 36.25 MTP. The 2.872x paired gain compares kernel on/off with MTP already enabled.
Additional information
Related: #218 (CUDA small batches), #225 (Metal single-vector work).
Measured base:
0324c66521960d67aa7da8687fb1453a79a6565c, research R4/S1. PR base:0781925904391351963d499cb32cd735849b06a5. Same kernel arithmetic; unused geometries/tuning switches removed. Performance has not been remeasured on the new base.Three ABBA quartets per cell, six observations per variant. Short prompts, up to 128 generated tokens; 4K context/slot, full Metal offload, Flash Attention, F16 KV, 16 CPU threads, batch/ubatch 512. Plain/MTP are separate randomized cells. MTP uses a separately grafted pinned Qwen3.8 head and one draft token; this patch does not provide MTP.
Requirements