Repository navigation
Conversation
|
Hi @Sube-py, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
Sube-py
force-pushed
the
fix-metal-stq-supports-op
branch
from
August 31, 2026 04:23
b92e1bc to
063195d
Compare
STQ1_0 (coming via ggml-org#22836) has no Metal kernel. Without this guard, ggml_metal_supports_mul_mat_op returns true for it, the scheduler places STQ1_0 weights in Metal buffers when -ngl > 0, and compute aborts with 'Asserting on type 43' / EXC_BAD_ACCESS in the pipeline lookup. Reserves the GGML_TYPE_STQ1_0 enum id (43) so it matches the value the STQ PR uses. Once ggml-org#22836 merges, both branches agree. Assisted-by: OpenAI
Sube-py
force-pushed
the
fix-metal-stq-supports-op
branch
from
September 17, 2026 04:04
063195d to
5873329
Compare
Sube-py
marked this pull request as ready for review
September 17, 2026 04:04
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Fix a hard crash when loading an STQ1_0 GGUF model on Metal with
-ngl > 0.ggml_metal_supports_mul_mat_oponly excludesGGML_TYPE_NVFP4, soGGML_TYPE_STQ1_0(added by #22836; CPU/ARM-NEON only, no Metal kernel) passes the check. The scheduler then places STQ1_0 weights in Metal buffers, and at compute time thekernel_mul_mv_stq1_0_*pipeline does not exist - the lookup path aborts withAsserting on type 43/ crashes withEXC_BAD_ACCESSinggml_metal_encoder_set_pipeline.This PR adds
STQ1_0to the exclusion, the same pattern #19769 used for NVFP4. It also reserves theGGML_TYPE_STQ1_0enum id (43) to match the value used by #22836, so the two PRs compose cleanly regardless of merge order. On this base (9723942) the type id does not exist yet, hence the 2-line enum addition.Note: with this fix,
-ngl > 0on an STQ1_0 model works but is much slower than-ngl 0(Apple M3: tg64 ~22 t/s vs ~80 t/s) since every STQ matmul crosses the CPU/GPU boundary. Users should still prefer-ngl 0until a Metal kernel exists. The fix only removes the hard crash / abort.Additional information
Reproduced on Apple M3, macOS 26.5:
llama-bench -ngl 99andllama-server -ngl 5both crash before the fix; both run after it (STQ1_0 tensors stay on the CPU backend, the rest of the graph still offloads).Regarding the "multiple backends" note: this PR touches only ggml.h (type id reservation) and Metal. No CPU/CUDA/Vulkan code is modified - the one-line Metal change exists because that is where the crash originates.
Tests:
llama-bench -ngl 99,llama-server -ngl 5,llama-completion -ngl 99with Hy-MT2-1.8B-1.25bit: crash before, run after (Apple M3)./build/bin/test-backend-ops test -b metal: passes after the change (3/3 backends OK)-ngl 0Requirements