Repository navigation
Conversation
This comment was marked as resolved.
This comment was marked as resolved.
|
This seems interesting, @khosravipasha something for you? Naming-wise it should probably be |
|
@sjl623 If you are doing benchmarks, you could also add |
Hi @CISC, thanks for taking a look! We've renamed the format to STQ1_0 according to your suggestion. Happy to hear further thoughts :) |
Good call, thanks. I've added Q1_0 to the benchmark table. STQ1_0 actually still comes out ahead on tg128, which was a bit unexpected given Q1_0 is binary and has lower bpw. Separately, per the Sherry paper, the 1.25-bit ternary setting also has better accuracy than the 1-bit binary baseline. |
|
Hi @ggerganov, could you take a look at this PR when you have time? Thanks! |
Compile |
Hi, thanks for the report! Quick check: did you download the .gguf from HF directly, or convert one yourself per the docs? The HF gguf was produced with an older quant type ID and isn't loadable as-is. Please re-convert following the doc instructions (we'll refresh the HF weights later). Also note: currently only ARM (Apple M-series) is supported; x86 isn't yet. |
I downloaded |
@cherish-ltt |
|
Using Android app demo from https://huggingface.co/tencent/Hy-MT1.5-1.8B-1.25bit-GGUF: |
@dima-xd , The source gguf of APK was replaced by mistake yesterday, now it should be ok. Just reinstall and redownload the gguf. |
Thanks, it works now! |
|
./build/bin/llama-completion still error |
Looks like it's being killed due to OOM. The default context (262144) makes the KV cache ~16 GiB, which may exceed host memory. Could you try again with |
./build/bin/llama-completion --model model_zoo/Hy-MT1.5-1.8B-STQ1_0.gguf -p "Translate the following segment into Vietnamese, without additional explanation. Hello " --jinja -ngl 0 -c 4096 -n 64 -st system_info: n_threads = 4 (n_threads_batch = 4) / 16 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 | sampler seed: 3567722247 Translate the following segment into Vietnamese, without additional explanation. Hello Xin chào.َايْتُمْ بَعْضُنُمْ بَعْلَيٰ ٌبَعْلَيٰ ٌبَعْلَيٰ… الْمُنْ common_perf_print: sampling time = 13.33 ms performance and result not good. it only run with arm?? |
Yes, currently the acceleration only supports ARM chips. You can try it on an M-series MacBook, or download our APK to experience it directly. |
|
approved. |
## Summary Adds two paths through the GGUF loader for GGML type 41 so mobius can ingest both standard and Tencent's custom variants of Q1_0. ### 1. Mainline Q1_0 — 1-bit binary Layout (18 bytes / 128 elements): `[fp16 d][16B packed 1-bit signs]`. Dequant: `bit ? +d : -d`. Used by [prism-ml/Bonsai](https://huggingface.co/collections/prism-ml/bonsai) checkpoints. Mapped to `MatMulNBits` `bits=2`, `zero_point=1`, codes ∈ {0, 2}, scale=d. ### 2. Tencent custom Q1_0 — 2-bit SEQ Used by [`AngelSlim/Hy-MT1.5-1.8B-2bit-GGUF`](https://huggingface.co/AngelSlim/Hy-MT1.5-1.8B-2bit-GGUF). Per-row layout: `[fp16 stored_scale][2-bit codes packed LSB-first]`, native block size 512, codebook `{−3, −1, +1, +3} · stored_scale`. Mainline llama.cpp **refuses to load** these files because every tensor after the first Q1_0 entry lands at the wrong offset. We bypass the size calculation by reading each tensor from its explicit GGUF offset. Two `MatMulNBits` representations, selected by a flag: | | default | `MOBIUS_TENCENT_Q1_0_USE_NATIVE_2BIT=1` | |---|---|---| | ORT bits | 4 (inflated `2c ∈ {0,2,4,6}`, integer zp=3) | 2 (codes pass-through, float zp=1.5) | | on-disk weight bpw | 4 (2× the source) | 2 (matches source) | | CPU decode throughput | **~32 tok/s** | **~0.24 tok/s** | | ORT version | any | **≥ 1.27** ([#28354](microsoft/onnxruntime#28354)) | Both representations produce **bit-identical dequantized weights** matching HF safetensors to bf16 rounding (`max_abs ≈ 2.5e-3` vs HF reference matmul). The 130× decode speed difference is purely the ORT CPU kernel: the bits=4 path is the mature MLAS fused path, the bits=2 + float-zp path is the kernel's own self-described "naive implementation, need to be optimized" scalar fallback. Tracked in [microsoft/onnxruntime#28552](microsoft/onnxruntime#28552); once that lands the default can flip. **Block-size workaround:** ORT silently returns zeros for `block_size > 256` ([#28551](microsoft/onnxruntime#28551)). Both representations therefore expose `block_size = 128` by replicating each native 512-element scale across 4 sub-blocks. ## Other infrastructure changes - New flag [`tencent_q1_0_use_native_2bit`](src/mobius/_flags.py) (default `False`). - `QuantizedLinear` accepts `bits ∈ {2, 4, 8}` and a new `zero_point_dtype` argument. `UINT8` (default) keeps the bit-packed integer form; float dtypes produce one un-packed value per block. - `QuantizationConfig` gains `float_zero_point`; `TextModel` passes `config.dtype` as the zp dtype when it is set. - GGUF → `ArchitectureConfig` derivation now sets `rope_type` from `<arch>.rope.scaling.type` (defaulting to `"default"` when `rope.freq_base` is present). Without this fix, models with `rope.scaling.type = "none"` were built with no RoPE at all. - For `hunyuan-dense` GGUFs whose `rope.freq_base` exceeds `1e6` (Tencent's pipeline bakes the dynamic-NTK exponent into a static value), the original HF config is restored: `rope_type="dynamic"`, `rope_theta=10000`, `alpha=1000`. Otherwise long prompts diverge. - New GGUF arch mapping `hunyuan-dense → hunyuan_v1_dense` with `attn_q_norm`/`attn_k_norm` tensor name mappings. - `_detect_quant_params` returns `block_size` explicitly so the `QuantizationConfig` matches the actual per-type group size. ## Verification End-to-end on `tencent/Hy-MT1.5-1.8B-2bit`: - **Per-tensor dequantization** matches HF safetensors to bf16-rounding precision (max abs `2.3e-5`). - **Single-layer ORT matmul** vs HF dequantized matmul: max abs `2.5e-3`, mean `2.8e-4`. - **Short-prompt greedy generation** with the chat template produces `Hello, world!<EOS>` token-for-token identical to HF. - **Throughput** with default flag: 0.13 s prefill, 32.8 tok/s decode on CPU EP. Tests: - Split into `TestTencentQ10DefaultInflated4Bit`, `TestTencentQ10NativeBits2`, and `TestTencentQ10Shared` covering both flag values + shape/error invariants. - Full test suite (`tests/build_graph_test.py src/`): **2752 passed, 0 failures**. ## Out of scope / follow-ups - 1.25-bit Sherry (`AngelSlim/Hy-MT1.5-1.8B-1.25bit-GGUF`, [llama.cpp PR #22836](ggml-org/llama.cpp#22836) `STQ1_0`): needs a ternary-with-structured-sparsity codebook that doesn't map cleanly to `MatMulNBits`. Tracked in [microsoft/onnxruntime#28549](microsoft/onnxruntime#28549). - ORT MLAS fast path for `bits=2` + float zp: [microsoft/onnxruntime#28552](microsoft/onnxruntime#28552). When that lands, flip the flag default and remove the bits=4 inflation path. - ORT silent zero output for `block_size > 256`: [microsoft/onnxruntime#28551](microsoft/onnxruntime#28551). --------- Signed-off-by: Justin Chu <justinchu@microsoft.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: Copilot Autofix powered by AI <223894421+github-code-quality[bot]@users.noreply.github.com>
|
@sjl623 你好,实测 D8300s上测试,STQ1_0 prefill 速度(41 t/s)仅为 Q4_K_M(79 t/s)的一半, decode 速度(27 t/s)超过 Q4_K_M(22 t/s),感觉相较于Q4 decode 的优势不明显,prefill的差距很大,有优化计划吗 |
|
This needs to be rebased and the ggml type id needs adjusting. |
STQ1_0 performs similarly to the existing ternary formats(TQ1_0, TQ2_0). Q4 benefits from mature tiled GEMM kernels, while ternary formats currently only have vec-dot kernels. Dedicated GEMM/GEMV optimizations may be explored in the future. |
Done. |
Squashed merge of PR ggml-org#22836 (head sjl623/llama.cpp@7e74b8296). Assisted-by: Claude Opus 4.8 (1M context)
|
~/llama.cpp$ cmake --build build --config Release complie with an error |
1.3125 bpw quantization. Each block of 256 elements stores 64 groups of 4
ternary lanes (-1/0/+1) with the constraint that every group has exactly
one zero and three non-zero lanes of identical magnitude. The ternary
pattern is encoded as a 4-bit codebook index plus a 1-bit global sign,
yielding 32 patterns over 4 lanes (5 bits / 4 lanes = 1.25 bpw payload),
plus a per-block fp16 scale (0.0625 bpw) for 1.3125 bpw total.
Components:
- block_stq_0 layout and codebook in ggml-common.h
- reference quantize/dequantize/validate in ggml-quants.{h,c}
- generic CPU vec_dot in ggml-cpu/quants.{h,c}
- ARM NEON vec_dot using vqtbl2q for codebook lookup, vdotq_s32 for
accumulation, plus vld4q-based in-place repack of Q8_K activations
- enum slots: GGML_TYPE_STQ_0 = 42, LLAMA_FTYPE_MOSTLY_STQ_0 = 41,
GGMLQuantizationType.STQ_0 = 42, LlamaFileType.MOSTLY_STQ_0 = 41
- llama-quantize CLI option "STQ_0"
x86 port of the ARM NEON STQ1_0 kernel from Tencent Hunyuan PR ggml-org#22836. AVX2 codebook lookup + maddubs, generic fallback when AVX2 is unavailable.
|
https://huggingface.co/AngelSlim/Hy4-preview-GGUF/tree/main New high profile stq1_0 quant. |
|
Hi @eauchs — great PR. We tested STQ1_0 locally (Hunyuan-MT 1.8B and a custom (Also noticed @Green-Sky's pointer to the new Hy4-preview-GGUF — its STQ1_0 1. x86 AVX2 vec_dot kernel (the merge blocker) — patch readyThe PR currently only ships ARM NEON + scalar fallback, so x86 builds fall back
2. Reference encoder is not SSE-optimal — patch ready
3. Pre-rename official GGUFs fail to load — enum mismatch (older files only)The GGUFs published at AngelSlim/Hy-MT1.5-1.8B-1.25bit-GGUF (April/May) encode all The new Hy4-preview-STQ1_0.gguf already uses 43 throughout, so this only affects Validation
Both patches are against 1e411d8 on this PR's branch — happy to split/rebase Patch 0001 — ggml-cpu: AVX2 vec_dot kernel for STQ1_0 (98 lines)From 994c9a086b58db739fbb7d2b24a5b776535fd498 Mon Sep 17 00:00:00 2001
From: ForiLusa <4963634@qq.com>
Date: Sun, 30 Aug 2026 02:03:13 +0800
Subject: [PATCH 1/2] ggml-cpu: AVX2 vec_dot kernel for STQ1_0
Native AVX2 implementation of ggml_vec_dot_stq1_0_q8_K for x86:
sign-expanded codebook select via pshufb+blendv, four 2-bit lane
planes matched against contiguous stride-16 q8 slices, (q-1)
correction from y block sums.
Fixes vs earlier WIP draft:
- half 1 groups (32..63) must OR sign bytes 4..7, not 0..3
- final reduction must cvtepi32_ps the int32 accumulator;
castsi256_ps reinterprets bit patterns and turns negative
partial sums into NaNs
Hunyuan-MT 1.8B STQ1_0: coherent output at ~33 t/s (16-thread
Zen4-ish desktop), PPL parity with q8_0 on the test set;
scalar fallback was 5.5 t/s.
---
ggml/src/ggml-cpu/arch-fallback.h | 1 -
ggml/src/ggml-cpu/arch/x86/quants.c | 98 +++++++++++++++++++++++++++++
2 files changed, 98 insertions(+), 1 deletion(-)
diff --git a/ggml/src/ggml-cpu/arch-fallback.h b/ggml/src/ggml-cpu/arch-fallback.h
index 259140f..eea39ec 100644
--- a/ggml/src/ggml-cpu/arch-fallback.h
+++ b/ggml/src/ggml-cpu/arch-fallback.h
@@ -85,7 +85,6 @@
#elif defined(__x86_64__) || defined(__i386__) || defined(_M_IX86) || defined(_M_X64)
// quants.c
#define ggml_vec_dot_q2_0_q8_0_generic ggml_vec_dot_q2_0_q8_0
-#define ggml_vec_dot_stq1_0_q8_K_generic ggml_vec_dot_stq1_0_q8_K
// repack.cpp
#define ggml_quantize_mat_q8_0_4x4_generic ggml_quantize_mat_q8_0_4x4
#define ggml_quantize_mat_q8_K_4x4_generic ggml_quantize_mat_q8_K_4x4
diff --git a/ggml/src/ggml-cpu/arch/x86/quants.c b/ggml/src/ggml-cpu/arch/x86/quants.c
index ea54cfe..5bc85b7 100644
--- a/ggml/src/ggml-cpu/arch/x86/quants.c
+++ b/ggml/src/ggml-cpu/arch/x86/quants.c
@@ -1763,6 +1763,104 @@ void ggml_vec_dot_q2_K_q8_K(int n, float * GGML_RESTRICT s, size_t bs, const voi
#endif
}
+void ggml_vec_dot_stq1_0_q8_K(int n, float * GGML_RESTRICT s, size_t bs, const void * GGML_RESTRICT vx, size_t bx, const void * GGML_RESTRICT vy, size_t by, int nrc) {
+ assert(nrc == 1);
+ UNUSED(nrc);
+ UNUSED(bx);
+ UNUSED(by);
+ UNUSED(bs);
+
+ const block_stq1_0 * GGML_RESTRICT x = vx;
+ const block_q8_K * GGML_RESTRICT y = vy;
+
+ const int nb = n / QK_K;
+
+#if defined(__AVX2__)
+ const __m128i m0f = _mm_set1_epi8(0x0F);
+ const __m128i m03 = _mm_set1_epi8(0x03);
+ const __m128i v16 = _mm_set1_epi8(16);
+ const __m128i cb_lo = _mm_loadu_si128((const __m128i *) stq1_0_codebook);
+ const __m128i cb_hi = _mm_loadu_si128((const __m128i *) (stq1_0_codebook + 16));
+ const __m256i ones16 = _mm256_set1_epi16(1);
+
+ float sumf = 0.0f;
+ for (int i = 0; i < nb; ++i) {
+ __m256i acc = _mm256_setzero_si256();
+
+ const uint8_t * sg = x[i].sign;
+ // expand the 8 sign bytes to 64 group sign bytes (0x10 / 0x00)
+ uint8_t sgn[64];
+ for (int j = 0; j < 64; ++j) {
+ sgn[j] = ((sg[j >> 3] >> (j & 7)) & 1) ? 0x10 : 0x00;
+ }
+
+ const __m128i qs0 = _mm_loadu_si128((const __m128i *) (x[i].qs));
+ const __m128i qs1 = _mm_loadu_si128((const __m128i *) (x[i].qs + 16));
+
+ // two 16-group chunks: chunk h covers groups i*32+16h..+15 and y bytes [i*128+64h .. +63]
+ for (int h = 0; h < 2; ++h) {
+ // sign select bytes for this half: groups 32h..32h+31 = sign bytes 4h..4h+3
+ const __m128i s0 = _mm_loadu_si128((const __m128i *) (sgn + 32*h));
+ const __m128i s1 = _mm_loadu_si128((const __m128i *) (sgn + 32*h + 16));
+ const __m128i packed = (h == 0) ? qs0 : qs1;
+ const __m128i lo = _mm_and_si128(packed, m0f);
+ const __m128i hi = _mm_and_si128(_mm_srli_epi16(packed, 4), m0f);
+ const __m128i ev = _mm_unpacklo_epi8(lo, hi); // groups h*32+0..7 (code+sign OR'd)
+ const __m128i od = _mm_unpackhi_epi8(lo, hi); // groups h*32+8..15
+
+ const __m128i evs = _mm_or_si128(ev, s0);
+ const __m128i ods = _mm_or_si128(od, s1);
+ // select the codebook half by the sign bit: 0x80 bytes take cb_hi, others cb_lo
+ const __m128i msk_e = _mm_slli_epi16(_mm_and_si128(evs, _mm_set1_epi8(0x10)), 3);
+ const __m128i msk_o = _mm_slli_epi16(_mm_and_si128(ods, _mm_set1_epi8(0x10)), 3);
+ const __m128i sel_e = _mm_blendv_epi8(_mm_shuffle_epi8(cb_lo, evs),
+ _mm_shuffle_epi8(cb_hi, _mm_sub_epi8(evs, v16)), msk_e);
+ const __m128i sel_o = _mm_blendv_epi8(_mm_shuffle_epi8(cb_lo, ods),
+ _mm_shuffle_epi8(cb_hi, _mm_sub_epi8(ods, v16)), msk_o);
+ const int8_t * q8 = y[i].qs + 128*h;
+ acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+ _mm256_cvtepi8_epi16(_mm_and_si128(sel_e, m03)),
+ _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 0)))));
+ acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+ _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_e, 2), m03)),
+ _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 16)))));
+ acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+ _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_e, 4), m03)),
+ _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 32)))));
+ acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+ _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_e, 6), m03)),
+ _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 48)))));
+ acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+ _mm256_cvtepi8_epi16(_mm_and_si128(sel_o, m03)),
+ _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 64)))));
+ acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+ _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_o, 2), m03)),
+ _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 80)))));
+ acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+ _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_o, 4), m03)),
+ _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 96)))));
+ acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+ _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_o, 6), m03)),
+ _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 112)))));
+ }
+
+ // (q-1) correction: subtract the y sum over the block's 256 positions
+ const __m256i ys = _mm256_loadu_si256((const __m256i *) y[i].bsums);
+ acc = _mm256_sub_epi32(acc, _mm256_madd_epi16(ys, _mm256_set1_epi16(1)));
+ // cvtepi32_ps converts the int32 lanes; castsi256_ps would reinterpret the
+ // bit patterns and turn negative partial sums into NaNs
+ sumf += (GGML_CPU_FP16_TO_FP32(x[i].d) * y[i].d) * hsum_float_8(_mm256_cvtepi32_ps(acc));
+ }
+
+ *s = sumf;
+#else
+ UNUSED(x);
+ UNUSED(y);
+ UNUSED(nb);
+ ggml_vec_dot_stq1_0_q8_K_generic(n, s, bs, vx, bx, vy, by, nrc);
+#endif
+}
+
void ggml_vec_dot_q3_K_q8_K(int n, float * GGML_RESTRICT s, size_t bs, const void * GGML_RESTRICT vx, size_t bx, const void * GGML_RESTRICT vy, size_t by, int nrc) {
assert(n % QK_K == 0);
assert(nrc == 1);
--
2.54.0.windows.1
Patch 0002 — quantize: SSE-optimal scale search for STQ1_0 (48 lines)From 01bad5d3247d3b04b6c467e340f95c707f6fccae Mon Sep 17 00:00:00 2001
From: ForiLusa <4963634@qq.com>
Date: Sun, 30 Aug 2026 11:55:29 +0800
Subject: [PATCH 2/2] quantize: SSE-optimal scale search for STQ1_0
With the per-group zero assignment fixed, the block reconstruction
error is sum(|x| - d)^2 over non-zero lanes, whose closed-form
minimizer is the mean |x| of those lanes - not amax. Grid-search
between the mean and amax per block and keep the lowest-SSE scale.
Format-compatible with existing kernels. Nemotron-H-8B f16->STQ1_0
round-trip RMSE 0.434 -> 0.261; wiki PPL (ffn-only ternary mix,
rest q8_0) 45307 -> 83.
---
ggml/src/ggml-quants.c | 50 ++++++++++++++++++++++++++++++++++++++++--
1 file changed, 48 insertions(+), 2 deletions(-)
diff --git a/ggml/src/ggml-quants.c b/ggml/src/ggml-quants.c
index 00ff915..d0c2ba7 100644
--- a/ggml/src/ggml-quants.c
+++ b/ggml/src/ggml-quants.c
@@ -2424,12 +2424,58 @@ void quantize_row_stq1_0_ref(const float * GGML_RESTRICT x, block_stq1_0 * GGML_
const float a = fabsf(x[j]);
if (a > amax) amax = a;
}
- y[i].d = GGML_FP32_TO_FP16(amax);
// STQ1_0 forces exactly one zero per group of 4. Groups are stride-16
// within each 64-weight chunk: group g (chunk-local in 0..15) covers
// {x[c*64 + g + p*16] : p in 0..3}. Pick the smallest-|x| lane as zero;
- // project the other 3 onto {-d, +d} via sign.
+ // project the other 3 onto {-d, +d}.
+ //
+ // d = amax is not SSE-optimal: with the zero assignment fixed, the
+ // block error is sum(|x| - d)^2 over non-zero lanes, whose closed-form
+ // minimizer is the mean |x| of those lanes. Grid-search between the
+ // mean and amax and keep the candidate with the lowest block SSE.
+ float mean_abs = 0.0f;
+ int n_nonzero = 0;
+ for (int g = 0; g < QK_K/4; ++g) {
+ const int chunk = g / 16;
+ const int gloc = g % 16;
+ const float * base = x + chunk*64 + gloc;
+ for (int p = 0; p < 4; ++p) {
+ mean_abs += fabsf(base[p*16]);
+ }
+ n_nonzero += 3;
+ }
+ mean_abs = (n_nonzero > 0) ? mean_abs / n_nonzero : amax;
+
+ float best_d = amax;
+ float best_sse = 1e30f;
+ for (int t = 0; t <= 7; ++t) {
+ const float d = mean_abs + (amax - mean_abs) * (t / 7.0f);
+ float sse = 0.0f;
+ for (int g = 0; g < QK_K/4; ++g) {
+ const int chunk = g / 16;
+ const int gloc = g % 16;
+ const float * base = x + chunk*64 + gloc;
+ int zero_pos = 0;
+ float min_abs = fabsf(base[0]);
+ for (int p = 1; p < 4; ++p) {
+ const float a = fabsf(base[p*16]);
+ if (a < min_abs) { min_abs = a; zero_pos = p; }
+ }
+ for (int p = 0; p < 4; ++p) {
+ if (p == zero_pos) {
+ sse += base[p*16] * base[p*16];
+ } else {
+ const float e = fabsf(base[p*16]) - d;
+ sse += e * e;
+ }
+ }
+ }
+ if (sse < best_sse) { best_sse = sse; best_d = d; }
+ }
+
+ y[i].d = GGML_FP32_TO_FP16(best_d);
+
for (int g = 0; g < QK_K/4; ++g) {
const int chunk = g / 16;
const int gloc = g % 16;
--
2.54.0.windows.1
|
STQ1_0 (coming via ggml-org#22836) has no Metal kernel. Without this guard, ggml_metal_supports_mul_mat_op returns true for it, the scheduler places STQ1_0 weights in Metal buffers when -ngl > 0, and compute aborts with 'Asserting on type 43' / EXC_BAD_ACCESS in the pipeline lookup. Reserves the GGML_TYPE_STQ1_0 enum id (43) so it matches the value the STQ PR uses. Once ggml-org#22836 merges, both branches agree.
|
The Hy4 merge made this more urgent: mainline still has no STQ1_0, so the AngelSlim Hy4 STQ1_0 GGUFs fail on stock builds ("invalid ggml type 43" — saw a user report exactly this a few days ago). The two patches I posted above (AVX2 vec_dot kernel + SSE scale search) still apply to the current head, re-checked today. Numbers from when I wrote them, Hunyuan-MT 1.8B, 8 threads on a 9800X3D: 5.5 t/s with the scalar fallback, 34.2 t/s with the kernel. That's Q8_0 territory at 1.31 bpw. I can keep the patches rebased on this branch and handle the x86 side in review, so missing x86 support isn't a reason to hold the merge. For anyone blocked on type 43 right now: it's only the type-field enum in the GGUFs (42 vs 43), rewriting those fields makes the model load fine. Small python script does it, ask if you need it. |
|
I re-measured on
The load failure has an exact cause. On master On the rebase question: the head I merged it in a worktree, built on an M1 Pro with Metal off, and ran
One thing that does not help. Master now enables |
|
@devYRPauli thanks for doing this legwork, it's more than I had. Good to know test-quantize-fns runs on the x86 runners with STQ1_0 bounds already registered, so the AVX2 kernel from my patch gets CI coverage as soon as it lands. Agreed on the backend-ops part. Adding STQ1_0 to those lists would just print Skipping CPU backend, since the CPU is the reference the others get compared against. Honest coverage for the CPU kernel today is the error-bounds test plus real-weights runs (that's where my 34 t/s number came from). If a proper cross-backend test shows up that treats the CPU as an implementation instead of the reference, I'll wire STQ1_0 into it. And if anyone hits the type 43 wall on the old AngelSlim files, ping me here, the fix is a small field-rewrite script, I'll paste it. |
STQ1_0 (coming via ggml-org#22836) has no Metal kernel. Without this guard, ggml_metal_supports_mul_mat_op returns true for it, the scheduler places STQ1_0 weights in Metal buffers when -ngl > 0, and compute aborts with 'Asserting on type 43' / EXC_BAD_ACCESS in the pipeline lookup. Reserves the GGML_TYPE_STQ1_0 enum id (43) so it matches the value the STQ PR uses. Once ggml-org#22836 merges, both branches agree. Assisted-by: OpenAI


Introduction
This PR adds STQ1_0 (Sparse Ternary Quantization), a new hardware-efficient ternary quantization kernel to support Sherry quantization, which recently accepted to ACL 2026 (Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification), where each weight is constrained to {-d, 0, +d} with the structural rule that exactly one of every four lanes is zero, yielding 1.3125 bits per weight (5 bits per 4-weight group, i.e. a 4-bit codebook index + 1-bit sign, plus a single
fp16scale per 256-weight block: 42 B / 256 = 1.3125 bpw) while admitting fast SIMD decode through a 32-entry codebook lookup.In llama.cpp terms, the 3:4 pattern lets STQ1_0 sit between the existing ternary formats: smaller than TQ2_0 (1.3125 vs 2.0625 bpw) and faster than TQ1_0, whose 1.6875-bit 3-way packing is SIMD-unfriendly; STQ1_0's power-of-two 4-way blocks decode directly through
vqtbl2q+vdotq_s32with no bit-shuffling.Below is a more concrete walkthrough of how our kernel actually decodes a block: the on-disk layout (

qs[32]4-bit codebook indices +sign[8]1-bit per group + anfp16scale), the codebook of 4-ternary lanes, and a step-by-step dequantization example showingcodebook lookup → sign flip → scale.Building on this method, Tencent Hunyuan has applied STQ1_0 to compress the Hy-MT1.5-1.8B model and released it to the open-source community, where it has received broad attention: the model has surpassed 16k+ downloads on Hugging Face within one week (4/29-5/5): AngelSlim/Hy-MT1.5-1.8B-1.25bit. In parallel, Tencent Hunyuan has built an on-device offline translation app on top of
llama.cpp; this PR contributes the 1.25-bit kernel implementation that powers it.There has also been community interest in adding Sherry support to llama.cpp from both directions: users have asked for this format on the llama.cpp side (discussion #19123) and on the model card side (Hy-MT1.5-1.8B-1.25bit / discussions / 3, specifically requesting the kernel). This PR is the upstream answer to those asks.
STQ1_0: Stride-16 Sparsity Layout
Notably, STQ1_0's 3:4 sparsity adopts a stride-16 grouping pattern, which is key to achieving efficient SIMD execution without extra data shuffling. Specifically, instead of grouping 4 consecutive weights (w0, w1, w2, w3), STQ1_0 groups weights that are stride-16 within each 64-weight chunk (e.g. w0, w16, w32, w48), where exactly one of the four is zero and the other three take values in {−d, +d}.
The motivation is SIMD alignment. After the codebook unpack (SHR 0/2/4/6 + mask), each NEON lane register (sqx0–sqx3) holds a contiguous run of weight indices. Since standard Q8_K activation quantization stores y values sequentially in memory, the lane contents and y are naturally aligned — a plain
vld1q_s8at offsets 0/16/32/48 is all that is needed, with no deinterleave or repack required.Acknowledge
Special thanks to @Little0o0 for proposing the Sherry method and for the helpful discussions on adapting its 3:4 packing layout to llama.cpp.
How to use it
1. Clone and build
cd llama.cpp cmake -B build cmake --build build --config Release -j2. Download the HF model
pip install huggingface_hub huggingface-cli download AngelSlim/Hy-MT1.5-1.8B-1.25bit \ --local-dir Hy-MT1.5-1.8B-1.25bitModel card: AngelSlim/Hy-MT1.5-1.8B-1.25bit.
3. Convert HF → GGUF (bf16)
python convert_hf_to_gguf.py Hy-MT1.5-1.8B-1.25bit \ --outfile Hy-MT1.5-1.8B-bf16.gguf \ --outtype bf164. Quantize bf16 → STQ1_0
./build/bin/llama-quantize \ Hy-MT1.5-1.8B-bf16.gguf \ Hy-MT1.5-1.8B-STQ1_0.gguf \ STQ1_05. Run a completion (CPU only, with chat template)
./build/bin/llama-completion \ -m Hy-MT1.5-1.8B-STQ1_0.gguf \ -ngl 0 --jinja \ -n 64 \ -p "translate to chinese: hello"Performance
Benchmarked on Apple M4 Pro (12 cores: 8 P + 4 E, 24 GB unified memory, macOS 26.3.1) with
-ngl 0so all layers run on CPU, exercising the ARM NEONvec_dotkernel introduced by this PR.To make the comparison against llama.cpp's existing ternary formats apples-to-apples, we used the community-reproduced Sherry-1B reference checkpoint MoraxGeo/Sherry-1B-1.25bit-per-channel, converted it to bf16 GGUF once, and then re-quantized the same weights into all three ternary formats (
STQ1_0/TQ1_0/TQ2_0) so that any difference below comes purely from the quantization format and its kernel../build/bin/llama-bench \ -m ./model_zoo/Sherry-1B-1.25bit-STQ1_0.gguf,\ ./model_zoo/Sherry-1B-1.25bit-TQ1_0.gguf,\ ./model_zoo/Sherry-1B-1.25bit-TQ2_0.gguf \ -ngl 0Based on the results above, STQ1_0 has the smallest footprint of the three (~11% smaller than TQ1_0, ~20% smaller than TQ2_0) and is faster than TQ1_0. Compared to the 1-bit binary
Q1_0baseline, STQ1_0 only trades a ~6% size increase but still has ~35% highertg128throughput. What's more, per the Sherry paper (Fig. 6), 1.25-bit ternary matches 1.67-bit ternary accuracy at 25% fewer bits, while the 1-bit binary configuration shows a ~3 pp accuracy gap.Requirements