Skip to content

metal : add PTQ1_0 mat-vec for two to four columns - #262

Merged
bri-prism merged 1 commit into
PrismML-Eng:prismfrom
jasontitus:downstream/metal-ptq1-pr
Sep 25, 2026
Merged

bri-prism merged 1 commit into
PrismML-Eng:prismfrom
jasontitus:downstream/metal-ptq1-pr

Conversation

@jasontitus

Copy link
Copy Markdown

Overview

Add an opt-in Metal PTQ1_0/F32 mat-vec for 2-4 columns, sharing decoded weights across columns and activations across four output rows. Enable with GGML_METAL_PTQ1_MULTICOL=1; default remains off. Existing single-vector and larger-batch paths are unchanged. Five files; no benchmark framework, MTP tooling or device-specific dispatch.

Single-user Bonsai 2 PTQ1_0 native generation tok/s:

Device Original plain Original MTP Optimized plain Optimized MTP MTP gain over optimized plain
M1 Ultra 26.93 10.91 27.26 37.18 +36.4%
M5 Max 30.19 13.01 29.98 39.48 +31.7%

For an M5 single user, candidate MTP is about 32% faster generation than candidate plain. M5 full-request throughput, including prompt/request overhead, is 27.96 plain and 36.25 MTP. The 2.872x paired gain compares kernel on/off with MTP already enabled.

Additional information

Related: #218 (CUDA small batches), #225 (Metal single-vector work).

Measured base: 0324c66521960d67aa7da8687fb1453a79a6565c, research R4/S1. PR base: 0781925904391351963d499cb32cd735849b06a5. Same kernel arithmetic; unused geometries/tuning switches removed. Performance has not been remeasured on the new base.

Three ABBA quartets per cell, six observations per variant. Short prompts, up to 128 generated tokens; 4K context/slot, full Metal offload, Flash Attention, F16 KV, 16 CPU threads, batch/ubatch 512. Plain/MTP are separate randomized cells. MTP uses a separately grafted pinned Qwen3.8 head and one draft token; this patch does not provide MTP.

  • Current-base M1: Release build passed; 172/172 CPU-reference backend cases and 42/42 numerical checks passed with selector off and on. Two-chunk WikiText-2 perplexity: 8.7487 off, 8.7486 on (context 512, batch/ubatch 4).
  • Measured PTQ1_0 server output comparisons matched; all 168 M5 plain/MTP comparisons matched. Logits are not bitwise equal (M5 max absolute difference 0.0012482, max NMSE 3.003e-10).
  • M1 completed standard llama-bench. M5 short-prefill quartets failed stability checks; no accepted M5 llama-bench result is claimed. Thermal samples were predominantly Heavy (M1 88.7%, M5 83.3%).
  • Five M5 output comparisons differed for older ternary PQ2_0 at concurrency 4, including baseline repeats. That format does not use this kernel.
  • Full CI and final reduced-patch M5 validation remain outstanding; this is for draft review.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. OpenAI Codex assisted with implementation, testing, benchmarking, analysis and this description. Contributor manual review and ownership are required before submission.

@bri-prism bri-prism left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Tested on an M5 Pro (Metal), merged onto current prism (bda59ea); the merge is clean. test-backend-ops MUL_MAT PTQ1_0 passes 172/172 with GGML_METAL_PTQ1_MULTICOL off and on. Output with the switch on matches off on Ternary Bonsai 2 27B PTQ1_0 (llama-perplexity, ubatch 4: mean KLD ~0, max 3.7e-5, same top token 100%).

llama-bench, 27B PTQ1_0, -fa 1, two rounds, t/s: pp2 10.6 -> 37.6, pp3 15.0 -> 34.0, pp4 19.3 -> 39.9; pp1 and tg32 unchanged (25.1 / 25.8), and with the switch off it matches the prism base exactly. Without this, a two-column verify on Metal runs at under half the single-token speed, which is why speculative decoding (DFlash2 #261, MTP) loses on Metal today. Given there's no regression on the off path, it may be worth making this the default once it's been checked on one more Apple generation. LGTM.

@bri-prism
bri-prism merged commit df7c49e into PrismML-Eng:prism Sep 25, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants