Skip to content

HIP: fix template skip for DKQ > 256 mfma kernels - #29559

Merged
IMbackK merged 1 commit into
ggml-org:masterfrom
IMbackK:cdna_dkq_256_fix
Sep 28, 2026
Merged

IMbackK merged 1 commit into
ggml-org:masterfrom
IMbackK:cdna_dkq_256_fix

Conversation

@IMbackK

@IMbackK IMbackK commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Overview

Skip condition for mfma templates dident exclude ncols1*ncols2 == 32, which is what is actually launched.

Requirements

@IMbackK
IMbackK requested a review from a team as a code owner September 28, 2026 07:29
@IMbackK

IMbackK commented Sep 28, 2026

Copy link
Copy Markdown
Contributor Author

fixes #29552 (comment)

cc @JohannesGaessler

@github-actions github-actions Bot added ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend labels Sep 28, 2026
@IMbackK
IMbackK merged commit c2a9e16 into ggml-org:master Sep 28, 2026
9 of 10 checks passed
pierreguillot pushed a commit to Ircam-Partiels/llama.cpp that referenced this pull request Oct 1, 2026
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
edwardyoon pushed a commit to edwardyoon/focus-llama that referenced this pull request Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants