Skip to content

metal: FWHT perf optimizations - #29602

Merged
taronaeo merged 1 commit into
ggml-org:masterfrom
PrismML-Eng:metal-fwht-tg-512
Sep 29, 2026
Merged

taronaeo merged 1 commit into
ggml-org:masterfrom
PrismML-Eng:metal-fwht-tg-512

Conversation

@bri-prism

Copy link
Copy Markdown
Contributor

Overview

FWHT perf optimizations

Additional information

$ git log --oneline -1
f9351fb09 metal: FWHT perf optimizations
$ build/bin/test-backend-ops test -b MTL0 -o MUL_MAT_HADAMARD
Backend 1/3: MTL0
  32/32 tests passed
  Backend MTL0: OK
$ master/build/bin/test-backend-ops perf -b MTL0 -o MUL_MAT_HADAMARD
MUL_MAT_HADAMARD(type_a=f32,type_b=f32,m=512,n=1,k=512,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):                303030 runs -     3.36 us/run - 524.29 kFLOP/run - 156.10 GFLOPS
MUL_MAT_HADAMARD(type_a=f32,type_b=f32,m=256,n=1,k=256,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):                466830 runs -     2.15 us/run - 131.07 kFLOP/run -  60.94 GFLOPS
MUL_MAT_HADAMARD(type_a=f32,type_b=f32,m=256,n=2048,k=256,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):              98845 runs -    10.14 us/run - 268.44 MFLOP/run -  26.46 TFLOPS
MUL_MAT_HADAMARD(type_a=f32,type_b=f32,m=512,n=2048,k=512,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):              23312 runs -    43.04 us/run -   1.07 GFLOP/run -  24.95 TFLOPS
$ pr/build/bin/test-backend-ops perf -b MTL0 -o MUL_MAT_HADAMARD
MUL_MAT_HADAMARD(type_a=f32,type_b=f32,m=512,n=1,k=512,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):                507780 runs -     1.99 us/run - 524.29 kFLOP/run - 263.49 GFLOPS
MUL_MAT_HADAMARD(type_a=f32,type_b=f32,m=256,n=1,k=256,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):                475020 runs -     2.13 us/run - 131.07 kFLOP/run -  61.53 GFLOPS
MUL_MAT_HADAMARD(type_a=f32,type_b=f32,m=256,n=2048,k=256,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):              98472 runs -    10.16 us/run - 268.44 MFLOP/run -  26.43 TFLOPS
MUL_MAT_HADAMARD(type_a=f32,type_b=f32,m=512,n=2048,k=512,bs=[1,1],nr=[1,1],per=[0,1,2,3],k_v=0,o=1,src_overlap=0,m_v=0,pad=0):              50008 runs -    20.03 us/run -   1.07 GFLOP/run -  53.60 TFLOPS

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. Claude Code was used to help develop and test the fix and to format this description to the PR template. I reviewed every line and take full responsibility for the changes.

Assisted-by: Claude Code
@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) labels Sep 28, 2026
@ggerganov
ggerganov marked this pull request as ready for review September 28, 2026 18:26
@ggerganov
ggerganov requested a review from a team as a code owner September 28, 2026 18:26
@ggerganov ggerganov added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 28, 2026
@taronaeo
taronaeo merged commit 18bbc46 into ggml-org:master Sep 29, 2026
23 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
Assisted-by: Claude Code
(cherry picked from commit 18bbc46)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants