Skip to content

ggml-metal : exclude STQ1_0 from Metal mul_mat supports_op - #28067

Open
Sube-py wants to merge 1 commit into
ggml-org:masterfrom
Sube-py:fix-metal-stq-supports-op
Open

Sube-py wants to merge 1 commit into
ggml-org:masterfrom
Sube-py:fix-metal-stq-supports-op

Conversation

@Sube-py

@Sube-py Sube-py commented Aug 31, 2026 •

Copy link
Copy Markdown

Overview

Fix a hard crash when loading an STQ1_0 GGUF model on Metal with -ngl > 0.

ggml_metal_supports_mul_mat_op only excludes GGML_TYPE_NVFP4, so GGML_TYPE_STQ1_0 (added by #22836; CPU/ARM-NEON only, no Metal kernel) passes the check. The scheduler then places STQ1_0 weights in Metal buffers, and at compute time the kernel_mul_mv_stq1_0_* pipeline does not exist - the lookup path aborts with Asserting on type 43 / crashes with EXC_BAD_ACCESS in ggml_metal_encoder_set_pipeline.

This PR adds STQ1_0 to the exclusion, the same pattern #19769 used for NVFP4. It also reserves the GGML_TYPE_STQ1_0 enum id (43) to match the value used by #22836, so the two PRs compose cleanly regardless of merge order. On this base (9723942) the type id does not exist yet, hence the 2-line enum addition.

Note: with this fix, -ngl > 0 on an STQ1_0 model works but is much slower than -ngl 0 (Apple M3: tg64 ~22 t/s vs ~80 t/s) since every STQ matmul crosses the CPU/GPU boundary. Users should still prefer -ngl 0 until a Metal kernel exists. The fix only removes the hard crash / abort.

Additional information

Reproduced on Apple M3, macOS 26.5: llama-bench -ngl 99 and llama-server -ngl 5 both crash before the fix; both run after it (STQ1_0 tensors stay on the CPU backend, the rest of the graph still offloads).

Regarding the "multiple backends" note: this PR touches only ggml.h (type id reservation) and Metal. No CPU/CUDA/Vulkan code is modified - the one-line Metal change exists because that is where the crash originates.

Tests:

  • llama-bench -ngl 99, llama-server -ngl 5, llama-completion -ngl 99 with Hy-MT2-1.8B-1.25bit: crash before, run after (Apple M3)
  • ./build/bin/test-backend-ops test -b metal: passes after the change (3/3 backends OK)
  • Translation output with the Hy-MT2-1.8B-1.25bit model unchanged vs -ngl 0

Requirements

@Sube-py
Sube-py requested review from a team, CISC and ggerganov as code owners August 31, 2026 04:06
@github-actions github-actions Bot added testing Everything test related examples ggml changes relating to the ggml tensor library for machine learning Apple Metal https://en.wikipedia.org/wiki/Metal_(API) conversion labels Aug 31, 2026
@ggml-gh-bot

ggml-gh-bot Bot commented Aug 31, 2026

Copy link
Copy Markdown

Hi @Sube-py, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • PR Template not respected: Please respect the template when creating a new pull request. Make sure to fill out all required sections.

  • Multiple backend changes in one PR: When adding support for a new model or feature, focus on CPU support only in the initial PR. Add support for other backends like CUDA in follow-up PRs. If you have a good reason to modify multiple backends in one PR, please explain it.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@ggml-gh-bot ggml-gh-bot Bot added the draft PR will be changed to draft by github-actions bot label Aug 31, 2026
@github-actions
github-actions Bot marked this pull request as draft August 31, 2026 04:10
@github-actions github-actions Bot removed the draft PR will be changed to draft by github-actions bot label Aug 31, 2026
@Sube-py
Sube-py force-pushed the fix-metal-stq-supports-op branch from b92e1bc to 063195d Compare August 31, 2026 04:23
@Sube-py Sube-py changed the title ggml-metal : fix EXC_BAD_ACCESS when loading STQ1_0 models with -ngl > 0 ggml-metal : exclude STQ1_0 from Metal mul_mat supports_op Aug 31, 2026
STQ1_0 (coming via ggml-org#22836) has no Metal kernel. Without this guard,
ggml_metal_supports_mul_mat_op returns true for it, the scheduler
places STQ1_0 weights in Metal buffers when -ngl > 0, and compute
aborts with 'Asserting on type 43' / EXC_BAD_ACCESS in the
pipeline lookup.

Reserves the GGML_TYPE_STQ1_0 enum id (43) so it matches the
value the STQ PR uses. Once ggml-org#22836 merges, both branches agree.

Assisted-by: OpenAI
@Sube-py
Sube-py force-pushed the fix-metal-stq-supports-op branch from 063195d to 5873329 Compare September 17, 2026 04:04
@Sube-py
Sube-py marked this pull request as ready for review September 17, 2026 04:04

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Apple Metal https://en.wikipedia.org/wiki/Metal_(API) conversion examples ggml changes relating to the ggml tensor library for machine learning testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants