Skip to content

opencl: add bin kernel kernel_gemm_noshuffle_q5_k_f32_32b_trans_ila_a8_bin, kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin - #29401

Merged
max-krasnyansky merged 2 commits into
ggml-org:masterfrom
qualcomm:q5_k-a8-dp4a-bin-kernel
Sep 25, 2026
Merged

max-krasnyansky merged 2 commits into
ggml-org:masterfrom
qualcomm:q5_k-a8-dp4a-bin-kernel

Conversation

@shaofeiqi

@shaofeiqi shaofeiqi commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Overview

Add optimized Q5_K non-MoE GEMM (DP4A and non-DP4A) bin kernels and a matching GEMV kernel for Adreno.

Requirements

@shaofeiqi
shaofeiqi requested a review from a team as a code owner September 24, 2026 21:14
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend labels Sep 24, 2026
@lhez

lhez commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

@max-krasnyansky please review/ack again when you get a chance.

@lhez lhez added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 25, 2026
@max-krasnyansky
max-krasnyansky merged commit a25c986 into ggml-org:master Sep 25, 2026
25 of 29 checks passed
feal87 added a commit to feal87/myllama.cpp that referenced this pull request Sep 25, 2026
Merge upstream commits:
  - llama: add llama_prec_policy + model-driven W4A4 path (ggml-org#24364)
  - llama: fix tensor split for fused qkv with uneven K/V head sizes (ggml-org#29294)
  - metal: split fa kernels into per-dtype libraries (ggml-org#29329)
  - metal: FWHT kernels for block widths above 512 (ggml-org#29095)
  - CUDA: fuse RMS_NORM + SCALE into one kernel (ggml-org#29393)
  - common: extract shared unicode path/string helpers (ggml-org#29415)
  - common,rpc: simplify fs_create_directory_with_parents() (ggml-org#29432)
  - rpc: include nb in the get_alloc_size cache key (ggml-org#29283)
  - [SYCL] support sparse FA (ggml-org#28796)
  - musa: fix PH1 operator failures and build issues (ggml-org#29193)
  - HIP: bump HIP_VERSION required for fp8 (ggml-org#29231)
  - opencl: add q5_k bin kernel (ggml-org#29401)
  - hexagon: add q5_k quant type support (ggml-org#29123)
  - hexagon: use DMA for contiguous dim1 CONCAT (ggml-org#29404)
  - mtmd: fix mel preprocessor in LFM2 audio (ggml-org#29403)
  - vulkan: fix legacy GLSLC without cooperativeMatrix (ggml-org#29409)
  - gguf-py: ByteLevel processing defaults bos/eos to False (ggml-org#29422)
  - gguf-py: TemplateProcessing has final word on add_special_token (ggml-org#29417)

Assisted-by: Pi
sky-mighty pushed a commit to sky-mighty/llama.cpp that referenced this pull request Sep 26, 2026
…a8_bin`, `kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin` (ggml-org#29401)

* opencl: add A8 Q5_K non-MoE non dp4a + dp4a binary kernel

* opencl: fix s transpose - s only transposed for bin kernels

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
wanghqc added a commit to qualcomm/llama.cpp that referenced this pull request Sep 29, 2026
ggml-opencl.cpp: this branch's layout kept, the three upstream OpenCL changes
(ggml-org#29401 q5_K bin kernels, ggml-org#29439 q8_0 dp4a bin kernel, ggml-org#29503 bin kernel
loading condition) replayed onto it. The q5_K bin layout and the q8_0 dp4a bin
GEMM are opt-in here (GGML_OPENCL_Q5_K_BIN=1, GGML_OPENCL_Q8_0_BIN_DP4A=1):
by default they would take the spec/MTP verify widths from the cooperative-K
and narrow kernels. q5_K bin is limited to 2-D weights, q8_0 bin to N > 16.

common/speculative, server-context: this branch's llama_batch implementation
kept (tree drafting puts several seq_ids on one token, which common_batch
cannot hold); a common_batch overload of common_speculative_process converts
for the new callers, the mtmd post-decode callback takes the new embd batch,
and ggml-org#28876, ggml-org#29556 and ggml-org#29648 are applied to server-context.
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
…a8_bin`, `kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin` (ggml-org#29401)

* opencl: add A8 Q5_K non-MoE non dp4a + dp4a binary kernel

* opencl: fix s transpose - s only transposed for bin kernels

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…a8_bin`, `kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin` (ggml-org#29401)

* opencl: add A8 Q5_K non-MoE non dp4a + dp4a binary kernel

* opencl: fix s transpose - s only transposed for bin kernels

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
…a8_bin`, `kernel_gemm_noshuffle_q5_k_q8_1_dp4a_ila_a8_bin` (ggml-org#29401)

* opencl: add A8 Q5_K non-MoE non dp4a + dp4a binary kernel

* opencl: fix s transpose - s only transposed for bin kernels

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
(cherry picked from commit a25c986)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. OpenCL Issues specific to the OpenCL backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants