Skip to content

ggml-cpu : add STQ1_0 ternary quantization with ARM NEON vec_dot kernel - #22836

Open
sjl623 wants to merge 3 commits into
ggml-org:masterfrom
sjl623:STQ_0
Open

sjl623 wants to merge 3 commits into
ggml-org:masterfrom
sjl623:STQ_0

Conversation

@sjl623

@sjl623 sjl623 commented May 8, 2026 •

Copy link
Copy Markdown

Introduction

This PR adds STQ1_0 (Sparse Ternary Quantization), a new hardware-efficient ternary quantization kernel to support Sherry quantization, which recently accepted to ACL 2026 (Sherry: Hardware-Efficient 1.25-Bit Ternary Quantization via Fine-grained Sparsification), where each weight is constrained to {-d, 0, +d} with the structural rule that exactly one of every four lanes is zero, yielding 1.3125 bits per weight (5 bits per 4-weight group, i.e. a 4-bit codebook index + 1-bit sign, plus a single fp16 scale per 256-weight block: 42 B / 256 = 1.3125 bpw) while admitting fast SIMD decode through a 32-entry codebook lookup.

image

In llama.cpp terms, the 3:4 pattern lets STQ1_0 sit between the existing ternary formats: smaller than TQ2_0 (1.3125 vs 2.0625 bpw) and faster than TQ1_0, whose 1.6875-bit 3-way packing is SIMD-unfriendly; STQ1_0's power-of-two 4-way blocks decode directly through vqtbl2q + vdotq_s32 with no bit-shuffling.

Below is a more concrete walkthrough of how our kernel actually decodes a block: the on-disk layout (qs[32] 4-bit codebook indices + sign[8] 1-bit per group + an fp16 scale), the codebook of 4-ternary lanes, and a step-by-step dequantization example showing codebook lookup → sign flip → scale.
Clipboard_Screenshot_1778238386

Building on this method, Tencent Hunyuan has applied STQ1_0 to compress the Hy-MT1.5-1.8B model and released it to the open-source community, where it has received broad attention: the model has surpassed 16k+ downloads on Hugging Face within one week (4/29-5/5): AngelSlim/Hy-MT1.5-1.8B-1.25bit. In parallel, Tencent Hunyuan has built an on-device offline translation app on top of llama.cpp; this PR contributes the 1.25-bit kernel implementation that powers it.

There has also been community interest in adding Sherry support to llama.cpp from both directions: users have asked for this format on the llama.cpp side (discussion #19123) and on the model card side (Hy-MT1.5-1.8B-1.25bit / discussions / 3, specifically requesting the kernel). This PR is the upstream answer to those asks.

STQ1_0: Stride-16 Sparsity Layout

Notably, STQ1_0's 3:4 sparsity adopts a stride-16 grouping pattern, which is key to achieving efficient SIMD execution without extra data shuffling. Specifically, instead of grouping 4 consecutive weights (w0, w1, w2, w3), STQ1_0 groups weights that are stride-16 within each 64-weight chunk (e.g. w0, w16, w32, w48), where exactly one of the four is zero and the other three take values in {−d, +d}.

The motivation is SIMD alignment. After the codebook unpack (SHR 0/2/4/6 + mask), each NEON lane register (sqx0–sqx3) holds a contiguous run of weight indices. Since standard Q8_K activation quantization stores y values sequentially in memory, the lane contents and y are naturally aligned — a plain vld1q_s8 at offsets 0/16/32/48 is all that is needed, with no deinterleave or repack required.

Untitled-2025-05-13-1959 excalidraw

Acknowledge

Special thanks to @Little0o0 for proposing the Sherry method and for the helpful discussions on adapting its 3:4 packing layout to llama.cpp.

How to use it

1. Clone and build

cd llama.cpp
cmake -B build
cmake --build build --config Release -j

2. Download the HF model

pip install huggingface_hub
huggingface-cli download AngelSlim/Hy-MT1.5-1.8B-1.25bit \
    --local-dir Hy-MT1.5-1.8B-1.25bit

Model card: AngelSlim/Hy-MT1.5-1.8B-1.25bit.

3. Convert HF → GGUF (bf16)

python convert_hf_to_gguf.py Hy-MT1.5-1.8B-1.25bit \
    --outfile Hy-MT1.5-1.8B-bf16.gguf \
    --outtype bf16

4. Quantize bf16 → STQ1_0

./build/bin/llama-quantize \
    Hy-MT1.5-1.8B-bf16.gguf \
    Hy-MT1.5-1.8B-STQ1_0.gguf \
    STQ1_0

5. Run a completion (CPU only, with chat template)

./build/bin/llama-completion \
    -m Hy-MT1.5-1.8B-STQ1_0.gguf \
    -ngl 0 --jinja \
    -n 64 \
    -p "translate to chinese: hello"

Performance

Benchmarked on Apple M4 Pro (12 cores: 8 P + 4 E, 24 GB unified memory, macOS 26.3.1) with -ngl 0 so all layers run on CPU, exercising the ARM NEON vec_dot kernel introduced by this PR.

To make the comparison against llama.cpp's existing ternary formats apples-to-apples, we used the community-reproduced Sherry-1B reference checkpoint MoraxGeo/Sherry-1B-1.25bit-per-channel, converted it to bf16 GGUF once, and then re-quantized the same weights into all three ternary formats (STQ1_0 / TQ1_0 / TQ2_0) so that any difference below comes purely from the quantization format and its kernel.

./build/bin/llama-bench \
    -m ./model_zoo/Sherry-1B-1.25bit-STQ1_0.gguf,\
./model_zoo/Sherry-1B-1.25bit-TQ1_0.gguf,\
./model_zoo/Sherry-1B-1.25bit-TQ2_0.gguf \
    -ngl 0
model size params threads test t/s
llama 1B STQ1_0 — 1.31 bpw 358.00 MiB 1.24 B 8 pp512 732.69 ± 20.00
llama 1B STQ1_0 — 1.31 bpw 358.00 MiB 1.24 B 8 tg128 147.47 ± 1.36
llama 1B Q1_0 336.25 MiB 1.24 B 8 pp512 768.47 ± 14.75
llama 1B Q1_0 336.25 MiB 1.24 B 8 tg128 109.62 ± 16.93
llama 1B TQ1_0 — 1.69 bpw 401.50 MiB 1.24 B 8 pp512 728.69 ± 19.88
llama 1B TQ1_0 — 1.69 bpw 401.50 MiB 1.24 B 8 tg128 138.87 ± 0.96
llama 1B TQ2_0 — 2.06 bpw 445.00 MiB 1.24 B 8 pp512 689.25 ± 16.61
llama 1B TQ2_0 — 2.06 bpw 445.00 MiB 1.24 B 8 tg128 175.06 ± 1.30

Based on the results above, STQ1_0 has the smallest footprint of the three (~11% smaller than TQ1_0, ~20% smaller than TQ2_0) and is faster than TQ1_0. Compared to the 1-bit binary Q1_0 baseline, STQ1_0 only trades a ~6% size increase but still has ~35% higher tg128 throughput. What's more, per the Sherry paper (Fig. 6), 1.25-bit ternary matches 1.67-bit ternary accuracy at 25% fewer bits, while the 1-bit binary configuration shows a ~3 pp accuracy gap.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: Yes — for navigating the llama.cpp codebase and for sentence polishing of this PR description.

@sjl623
sjl623 requested review from CISC and ggerganov as code owners May 8, 2026 11:07
@ggml-gh-bot

This comment was marked as resolved.

@CISC

CISC commented May 8, 2026

Copy link
Copy Markdown
Member

This seems interesting, @khosravipasha something for you?

Naming-wise it should probably be STQ1_0.

@github-actions github-actions Bot added examples python python script changes ggml changes relating to the ggml tensor library for machine learning labels May 8, 2026
@Green-Sky

Green-Sky commented May 8, 2026 •

Copy link
Copy Markdown
Collaborator

@sjl623 If you are doing benchmarks, you could also add Q1_0 to the list, since it is another close quant with 1.125 bits-per-weight.

@sjl623 sjl623 changed the title ggml-cpu : add STQ_0 ternary quantization with ARM NEON vec_dot kernel ggml-cpu : add STQ1_0 ternary quantization with ARM NEON vec_dot kernel May 8, 2026
@sjl623

sjl623 commented May 8, 2026

Copy link
Copy Markdown
Author

This seems interesting, @khosravipasha something for you?

Naming-wise it should probably be STQ1_0.

Hi @CISC, thanks for taking a look! We've renamed the format to STQ1_0 according to your suggestion. Happy to hear further thoughts :)

@sjl623

sjl623 commented May 8, 2026

Copy link
Copy Markdown
Author

@sjl623 If you are doing benchmarks, you could also add Q1_0 to the list, since it is another close quant with 1.125 bits-per-weight.

Good call, thanks. I've added Q1_0 to the benchmark table. STQ1_0 actually still comes out ahead on tg128, which was a bit unexpected given Q1_0 is binary and has lower bpw. Separately, per the Sherry paper, the 1.25-bit ternary setting also has better accuracy than the 1-bit binary baseline.

@sjl623

sjl623 commented May 8, 2026 •

Copy link
Copy Markdown
Author

Hi @ggerganov, could you take a look at this PR when you have time? Thanks!

@cherish-ltt

Copy link
Copy Markdown
sudo ./build/bin/llama-completion --model /home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf -p "Translate the following segment into Chinese, without additional explanation:Hello" --jinja -ngl 0 -n 64 -st


warning: no usable GPU found, --gpu-layers option will be ignored
warning: one possible reason is that llama.cpp was compiled without GPU support
warning: consult docs/build.md for compilation instructions
main: llama backend init
main: load the model and apply lora adapter, if any
common_init_result: fitting params to device memory, for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on
common_params_fit_impl: getting device memory data for initial parameters:
gguf_init_from_file_ptr: tensor 'blk.0.attn_k_norm.weight' has offset 203154464, expected 203572256
gguf_init_from_file_ptr: failed to read tensor data
llama_model_load: error loading model: llama_model_loader: failed to load model from /home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf
llama_model_load_from_file_impl: failed to load model
common_fit_params: encountered an error while trying to fit params to free device memory: failed to load model
common_fit_params: fitting params to free memory took 0.08 seconds
gguf_init_from_file_ptr: tensor 'blk.0.attn_k_norm.weight' has offset 203154464, expected 203572256
gguf_init_from_file_ptr: failed to read tensor data
llama_model_load: error loading model: llama_model_loader: failed to load model from /home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf
llama_model_load_from_file_impl: failed to load model
common_init_from_params: failed to load model '/home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf'
main: error: unable to create context

Compile llama.cpp according to doc, run it with an error, hasn't the model been updated yet

@sjl623

sjl623 commented May 9, 2026

Copy link
Copy Markdown
Author
sudo ./build/bin/llama-completion --model /home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf -p "Translate the following segment into Chinese, without additional explanation:Hello" --jinja -ngl 0 -n 64 -st


warning: no usable GPU found, --gpu-layers option will be ignored
warning: one possible reason is that llama.cpp was compiled without GPU support
warning: consult docs/build.md for compilation instructions
main: llama backend init
main: load the model and apply lora adapter, if any
common_init_result: fitting params to device memory, for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on
common_params_fit_impl: getting device memory data for initial parameters:
gguf_init_from_file_ptr: tensor 'blk.0.attn_k_norm.weight' has offset 203154464, expected 203572256
gguf_init_from_file_ptr: failed to read tensor data
llama_model_load: error loading model: llama_model_loader: failed to load model from /home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf
llama_model_load_from_file_impl: failed to load model
common_fit_params: encountered an error while trying to fit params to free device memory: failed to load model
common_fit_params: fitting params to free memory took 0.08 seconds
gguf_init_from_file_ptr: tensor 'blk.0.attn_k_norm.weight' has offset 203154464, expected 203572256
gguf_init_from_file_ptr: failed to read tensor data
llama_model_load: error loading model: llama_model_loader: failed to load model from /home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf
llama_model_load_from_file_impl: failed to load model
common_init_from_params: failed to load model '/home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf'
main: error: unable to create context

Compile llama.cpp according to doc, run it with an error, hasn't the model been updated yet

Hi, thanks for the report! Quick check: did you download the .gguf from HF directly, or convert one yourself per the docs?

The HF gguf was produced with an older quant type ID and isn't loadable as-is. Please re-convert following the doc instructions (we'll refresh the HF weights later).

Also note: currently only ARM (Apple M-series) is supported; x86 isn't yet.

@cherish-ltt

Copy link
Copy Markdown
sudo ./build/bin/llama-completion --model /home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf -p "Translate the following segment into Chinese, without additional explanation:Hello" --jinja -ngl 0 -n 64 -st


warning: no usable GPU found, --gpu-layers option will be ignored
warning: one possible reason is that llama.cpp was compiled without GPU support
warning: consult docs/build.md for compilation instructions
main: llama backend init
main: load the model and apply lora adapter, if any
common_init_result: fitting params to device memory, for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on
common_params_fit_impl: getting device memory data for initial parameters:
gguf_init_from_file_ptr: tensor 'blk.0.attn_k_norm.weight' has offset 203154464, expected 203572256
gguf_init_from_file_ptr: failed to read tensor data
llama_model_load: error loading model: llama_model_loader: failed to load model from /home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf
llama_model_load_from_file_impl: failed to load model
common_fit_params: encountered an error while trying to fit params to free device memory: failed to load model
common_fit_params: fitting params to free memory took 0.08 seconds
gguf_init_from_file_ptr: tensor 'blk.0.attn_k_norm.weight' has offset 203154464, expected 203572256
gguf_init_from_file_ptr: failed to read tensor data
llama_model_load: error loading model: llama_model_loader: failed to load model from /home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf
llama_model_load_from_file_impl: failed to load model
common_init_from_params: failed to load model '/home/xxx/xxx/Hy-MT1.5-1.8B-1.25bit.gguf'
main: error: unable to create context

Compile llama.cpp according to doc, run it with an error, hasn't the model been updated yet按照文档编译 llama.cpp,运行时有错误,模型还没更新吗

Hi, thanks for the report! Quick check: did you download the .gguf from HF directly, or convert one yourself per the docs?你好,谢谢你的报告!快速确认一下:你是直接从 HF 下载的 .gguf 文件,还是根据文档自己转换的?

The HF gguf was produced with an older quant type ID and isn't loadable as-is. Please re-convert following the doc instructions (we'll refresh the HF weights later).HF gguf 是用较早的量化类型 ID 生产的,不能直接装填。请按照文档说明重新转换(我们稍后会刷新 HF 权重)。

Also note: currently only ARM (Apple M-series) is supported; x86 isn't yet.另外注意:目前仅支持 ARM(Apple M 系列);x86 还没有。

I downloaded gguf directly from HF, Llama.cpp was compiled by myself, and my platform is x86
Thank you for your work

@Little0o0

Little0o0 commented May 9, 2026 •

Copy link
Copy Markdown
Contributor

I downloaded gguf directly from HF, Llama.cpp was compiled by myself, and my platform is x86 Thank you for your work

@cherish-ltt gguf file on HF did not update yesterday. It is updated now. And I think STQ1_0 do not support x86 currently ?

@dima-xd

dima-xd commented May 10, 2026

Copy link
Copy Markdown

Using Android app demo from https://huggingface.co/tencent/Hy-MT1.5-1.8B-1.25bit-GGUF:

15:07:56.537  E  gguf_init_from_file_impl: tensor 'blk.0.attn_k.weight' has invalid ggml type 42. should be in [0, 42)
15:07:56.537  E  gguf_init_from_file_impl: failed to read tensor info
15:07:56.551  E  llama_model_load: error loading model: llama_model_loader: failed to load model from /data/user/0/com.tencent.hunyuan.angelslim/files/models/HY1.8B-MT-1.25bit.gguf
15:07:56.551  E  llama_model_load_from_file_impl: failed to load model
15:07:56.551  E  Error loading model
                 /data/user/0/com.tencent.hunyuan.angelslim/files/models/HY1.8B-MT-1.25bit.gguf
                 com.arm.aichat.UnsupportedArchitectureException

@Little0o0

Little0o0 commented May 10, 2026 •

Copy link
Copy Markdown
Contributor

Using Android app demo from https://huggingface.co/tencent/Hy-MT1.5-1.8B-1.25bit-GGUF:

15:07:56.537  E  gguf_init_from_file_impl: tensor 'blk.0.attn_k.weight' has invalid ggml type 42. should be in [0, 42)
15:07:56.537  E  gguf_init_from_file_impl: failed to read tensor info
15:07:56.551  E  llama_model_load: error loading model: llama_model_loader: failed to load model from /data/user/0/com.tencent.hunyuan.angelslim/files/models/HY1.8B-MT-1.25bit.gguf
15:07:56.551  E  llama_model_load_from_file_impl: failed to load model
15:07:56.551  E  Error loading model
                 /data/user/0/com.tencent.hunyuan.angelslim/files/models/HY1.8B-MT-1.25bit.gguf
                 com.arm.aichat.UnsupportedArchitectureException

@dima-xd , The source gguf of APK was replaced by mistake yesterday, now it should be ok. Just reinstall and redownload the gguf.

@dima-xd

dima-xd commented May 10, 2026

Copy link
Copy Markdown

Using Android app demo from https://huggingface.co/tencent/Hy-MT1.5-1.8B-1.25bit-GGUF:

15:07:56.537  E  gguf_init_from_file_impl: tensor 'blk.0.attn_k.weight' has invalid ggml type 42. should be in [0, 42)
15:07:56.537  E  gguf_init_from_file_impl: failed to read tensor info
15:07:56.551  E  llama_model_load: error loading model: llama_model_loader: failed to load model from /data/user/0/com.tencent.hunyuan.angelslim/files/models/HY1.8B-MT-1.25bit.gguf
15:07:56.551  E  llama_model_load_from_file_impl: failed to load model
15:07:56.551  E  Error loading model
                 /data/user/0/com.tencent.hunyuan.angelslim/files/models/HY1.8B-MT-1.25bit.gguf
                 com.arm.aichat.UnsupportedArchitectureException

@dima-xd , The source gguf of APK is replaced by mistake, now it should be ok.

Thanks, it works now!

@tritueviet

Copy link
Copy Markdown

./build/bin/llama-completion
--model ../model_zoo/Hy-MT1.5-1.8B-1.25bit-GGUF/Hy-MT1.5-1.8B-1.25bit.gguf
-p "Translate the following segment into Chinese, without additional explanation:Hello"
--jinja
-ngl 0
-n 64 -st
warning: no usable GPU found, --gpu-layers option will be ignored
warning: one possible reason is that llama.cpp was compiled without GPU support
warning: consult docs/build.md for compilation instructions
main: llama backend init
main: load the model and apply lora adapter, if any
common_init_result: fitting params to device memory, for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on
common_params_fit_impl: getting device memory data for initial parameters:
common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted |
common_memory_breakdown_print: | - Host | 17607 = 435 + 16384 + 788 |
common_params_fit_impl: projected to use 17607 MiB of host memory vs. 15908 MiB of total host memory
common_params_fit_impl: cannot meet free memory target of 1024 MiB, need to reduce device memory by 2723 MiB
common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted |
common_memory_breakdown_print: | - Host | 939 = 435 + 256 + 247 |
common_params_fit_impl: context size reduced from 262144 to 219904 -> need 2728 MiB less memory in total
common_params_fit_impl: entire model can be fit by reducing context
common_fit_params: successfully fit params to free device memory
common_fit_params: fitting params to free memory took 0.55 seconds
llama_model_loader: loaded meta data with 40 key-value pairs and 354 tensors from ../model_zoo/Hy-MT1.5-1.8B-1.25bit-GGUF/Hy-MT1.5-1.8B-1.25bit.gguf (version GGUF V3 (latest))
llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
llama_model_loader: - kv 0: general.architecture str = hunyuan-dense
llama_model_loader: - kv 1: general.type str = model
llama_model_loader: - kv 2: general.sampling.top_k i32 = 20
llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.800000
llama_model_loader: - kv 4: general.sampling.temp f32 = 0.700000
llama_model_loader: - kv 5: general.name str = Hy MT1.5 1.8B 1.25bit
llama_model_loader: - kv 6: general.finetune str = 1.25bit
llama_model_loader: - kv 7: general.basename str = Hy-MT1.5
llama_model_loader: - kv 8: general.size_label str = 1.8B
llama_model_loader: - kv 9: general.base_model.count u32 = 1
llama_model_loader: - kv 10: general.base_model.0.name str = HY MT1.5 1.8B
llama_model_loader: - kv 11: general.base_model.0.organization str = Tencent
llama_model_loader: - kv 12: general.base_model.0.repo_url str = https://huggingface.co/tencent/HY-MT1...
llama_model_loader: - kv 13: general.tags arr[str,5] = ["translation", "hy-mt", "quant", "1....
llama_model_loader: - kv 14: general.languages arr[str,1] = ["multilingual"]
llama_model_loader: - kv 15: hunyuan-dense.block_count u32 = 32
llama_model_loader: - kv 16: hunyuan-dense.context_length u32 = 262144
llama_model_loader: - kv 17: hunyuan-dense.embedding_length u32 = 2048
llama_model_loader: - kv 18: hunyuan-dense.feed_forward_length u32 = 6144
llama_model_loader: - kv 19: hunyuan-dense.attention.head_count u32 = 16
llama_model_loader: - kv 20: hunyuan-dense.attention.head_count_kv u32 = 4
llama_model_loader: - kv 21: hunyuan-dense.rope.freq_base f32 = 11158840.000000
llama_model_loader: - kv 22: hunyuan-dense.attention.layer_norm_rms_epsilon f32 = 0.000010
llama_model_loader: - kv 23: hunyuan-dense.attention.key_length u32 = 128
llama_model_loader: - kv 24: hunyuan-dense.attention.value_length u32 = 128
llama_model_loader: - kv 25: hunyuan-dense.rope.scaling.type str = none
llama_model_loader: - kv 26: hunyuan-dense.rope.scaling.factor f32 = 1.000000
llama_model_loader: - kv 27: hunyuan-dense.rope.scaling.original_context_length u32 = 262144
llama_model_loader: - kv 28: tokenizer.ggml.model str = gpt2
llama_model_loader: - kv 29: tokenizer.ggml.pre str = hunyuan-dense
llama_model_loader: - kv 30: tokenizer.ggml.tokens arr[str,120818] = ["!", """, "#", "$", "%", "&", "'", ...
llama_model_loader: - kv 31: tokenizer.ggml.token_type arr[i32,120818] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
llama_model_loader: - kv 32: tokenizer.ggml.merges arr[str,119758] = ["Ġ Ġ", "Ġ t", "Ġ a", "i n", "h e...
llama_model_loader: - kv 33: tokenizer.ggml.bos_token_id u32 = 120000
llama_model_loader: - kv 34: tokenizer.ggml.eos_token_id u32 = 120020
llama_model_loader: - kv 35: tokenizer.ggml.padding_token_id u32 = 120002
llama_model_loader: - kv 36: tokenizer.ggml.seperator_token_id u32 = 120007
llama_model_loader: - kv 37: tokenizer.chat_template str = {% if messages[0]['role'] == 'system'...
llama_model_loader: - kv 38: general.quantization_version u32 = 2
llama_model_loader: - kv 39: general.file_type u32 = 41
llama_model_loader: - type f32: 129 tensors
llama_model_loader: - type q6_K: 1 tensors
llama_model_loader: - type stq1_0: 224 tensors
print_info: file format = GGUF V3 (latest)
print_info: file type = STQ1_0 - 1.31 bpw ternary
print_info: file size = 435.61 MiB (2.04 BPW)
load: 0 unused tokens
load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect
load: printing all EOG tokens:
load: - 120020 ('<|hy_place▁holder▁no▁2|>')
load: special tokens cache size = 818
load: token to piece cache size = 0.8089 MB
print_info: arch = hunyuan-dense
print_info: vocab_only = 0
print_info: no_alloc = 0
print_info: n_ctx_train = 262144
print_info: n_embd = 2048
print_info: n_embd_inp = 2048
print_info: n_layer = 32
print_info: n_head = 16
print_info: n_head_kv = 4
print_info: n_rot = 128
print_info: n_swa = 0
print_info: is_swa_any = 0
print_info: n_embd_head_k = 128
print_info: n_embd_head_v = 128
print_info: n_gqa = 4
print_info: n_embd_k_gqa = 512
print_info: n_embd_v_gqa = 512
print_info: f_norm_eps = 0.0e+00
print_info: f_norm_rms_eps = 1.0e-05
print_info: f_clamp_kqv = 0.0e+00
print_info: f_max_alibi_bias = 0.0e+00
print_info: f_logit_scale = 0.0e+00
print_info: f_attn_scale = 0.0e+00
print_info: f_attn_value_scale = 0.0000
print_info: n_ff = 6144
print_info: n_expert = 0
print_info: n_expert_used = 0
print_info: n_expert_groups = 0
print_info: n_group_used = 0
print_info: causal attn = 1
print_info: pooling type = -1
print_info: rope type = 2
print_info: rope scaling = none
print_info: freq_base_train = 11158840.0
print_info: freq_scale_train = 1
print_info: n_ctx_orig_yarn = 262144
print_info: rope_yarn_log_mul = 0.0000
print_info: rope_finetuned = unknown
print_info: model type = 1.8B
print_info: model params = 1.79 B
print_info: general.name = Hy MT1.5 1.8B 1.25bit
print_info: vocab type = BPE
print_info: n_vocab = 120818
print_info: n_merges = 119758
print_info: BOS token = 120000 '<|hy_begin▁of▁sentence|>'
print_info: EOS token = 120020 '<|hy_place▁holder▁no▁2|>'
print_info: SEP token = 120007 '<|hy_Assistant|>'
print_info: PAD token = 120002 '<|hy_▁pad▁|>'
print_info: LF token = 185 'Ċ'
print_info: EOG token = 120020 '<|hy_place▁holder▁no▁2|>'
print_info: max token length = 1024
load_tensors: loading model tensors, this can take a while... (mmap = true, direct_io = false)
load_tensors: CPU_Mapped model buffer size = 435.61 MiB
.........................................................
common_init_result: added <|hy_place▁holder▁no▁2|> logit bias = -inf
llama_context: constructing llama_context
llama_context: n_seq_max = 1
llama_context: n_ctx = 219904
llama_context: n_ctx_seq = 219904
llama_context: n_batch = 2048
llama_context: n_ubatch = 512
llama_context: causal_attn = 1
llama_context: flash_attn = auto
llama_context: kv_unified = false
llama_context: freq_base = 11158840.0
llama_context: freq_scale = 1
llama_context: n_ctx_seq (219904) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
llama_context: CPU output buffer size = 0.46 MiB
llama_kv_cache: CPU KV buffer size = 13744.00 MiB
Killed

still error

@sjl623

sjl623 commented May 11, 2026

Copy link
Copy Markdown
Author

./build/bin/llama-completion --model ../model_zoo/Hy-MT1.5-1.8B-1.25bit-GGUF/Hy-MT1.5-1.8B-1.25bit.gguf -p "Translate the following segment into Chinese, without additional explanation:Hello" --jinja -ngl 0 -n 64 -st warning: no usable GPU found, --gpu-layers option will be ignored warning: one possible reason is that llama.cpp was compiled without GPU support warning: consult docs/build.md for compilation instructions main: llama backend init main: load the model and apply lora adapter, if any common_init_result: fitting params to device memory, for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on common_params_fit_impl: getting device memory data for initial parameters: common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | common_memory_breakdown_print: | - Host | 17607 = 435 + 16384 + 788 | common_params_fit_impl: projected to use 17607 MiB of host memory vs. 15908 MiB of total host memory common_params_fit_impl: cannot meet free memory target of 1024 MiB, need to reduce device memory by 2723 MiB common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | common_memory_breakdown_print: | - Host | 939 = 435 + 256 + 247 | common_params_fit_impl: context size reduced from 262144 to 219904 -> need 2728 MiB less memory in total common_params_fit_impl: entire model can be fit by reducing context common_fit_params: successfully fit params to free device memory common_fit_params: fitting params to free memory took 0.55 seconds llama_model_loader: loaded meta data with 40 key-value pairs and 354 tensors from ../model_zoo/Hy-MT1.5-1.8B-1.25bit-GGUF/Hy-MT1.5-1.8B-1.25bit.gguf (version GGUF V3 (latest)) llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output. llama_model_loader: - kv 0: general.architecture str = hunyuan-dense llama_model_loader: - kv 1: general.type str = model llama_model_loader: - kv 2: general.sampling.top_k i32 = 20 llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.800000 llama_model_loader: - kv 4: general.sampling.temp f32 = 0.700000 llama_model_loader: - kv 5: general.name str = Hy MT1.5 1.8B 1.25bit llama_model_loader: - kv 6: general.finetune str = 1.25bit llama_model_loader: - kv 7: general.basename str = Hy-MT1.5 llama_model_loader: - kv 8: general.size_label str = 1.8B llama_model_loader: - kv 9: general.base_model.count u32 = 1 llama_model_loader: - kv 10: general.base_model.0.name str = HY MT1.5 1.8B llama_model_loader: - kv 11: general.base_model.0.organization str = Tencent llama_model_loader: - kv 12: general.base_model.0.repo_url str = https://huggingface.co/tencent/HY-MT1... llama_model_loader: - kv 13: general.tags arr[str,5] = ["translation", "hy-mt", "quant", "1.... llama_model_loader: - kv 14: general.languages arr[str,1] = ["multilingual"] llama_model_loader: - kv 15: hunyuan-dense.block_count u32 = 32 llama_model_loader: - kv 16: hunyuan-dense.context_length u32 = 262144 llama_model_loader: - kv 17: hunyuan-dense.embedding_length u32 = 2048 llama_model_loader: - kv 18: hunyuan-dense.feed_forward_length u32 = 6144 llama_model_loader: - kv 19: hunyuan-dense.attention.head_count u32 = 16 llama_model_loader: - kv 20: hunyuan-dense.attention.head_count_kv u32 = 4 llama_model_loader: - kv 21: hunyuan-dense.rope.freq_base f32 = 11158840.000000 llama_model_loader: - kv 22: hunyuan-dense.attention.layer_norm_rms_epsilon f32 = 0.000010 llama_model_loader: - kv 23: hunyuan-dense.attention.key_length u32 = 128 llama_model_loader: - kv 24: hunyuan-dense.attention.value_length u32 = 128 llama_model_loader: - kv 25: hunyuan-dense.rope.scaling.type str = none llama_model_loader: - kv 26: hunyuan-dense.rope.scaling.factor f32 = 1.000000 llama_model_loader: - kv 27: hunyuan-dense.rope.scaling.original_context_length u32 = 262144 llama_model_loader: - kv 28: tokenizer.ggml.model str = gpt2 llama_model_loader: - kv 29: tokenizer.ggml.pre str = hunyuan-dense llama_model_loader: - kv 30: tokenizer.ggml.tokens arr[str,120818] = ["!", """, "#", "$", "%", "&", "'", ... llama_model_loader: - kv 31: tokenizer.ggml.token_type arr[i32,120818] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ... llama_model_loader: - kv 32: tokenizer.ggml.merges arr[str,119758] = ["Ġ Ġ", "Ġ t", "Ġ a", "i n", "h e... llama_model_loader: - kv 33: tokenizer.ggml.bos_token_id u32 = 120000 llama_model_loader: - kv 34: tokenizer.ggml.eos_token_id u32 = 120020 llama_model_loader: - kv 35: tokenizer.ggml.padding_token_id u32 = 120002 llama_model_loader: - kv 36: tokenizer.ggml.seperator_token_id u32 = 120007 llama_model_loader: - kv 37: tokenizer.chat_template str = {% if messages[0]['role'] == 'system'... llama_model_loader: - kv 38: general.quantization_version u32 = 2 llama_model_loader: - kv 39: general.file_type u32 = 41 llama_model_loader: - type f32: 129 tensors llama_model_loader: - type q6_K: 1 tensors llama_model_loader: - type stq1_0: 224 tensors print_info: file format = GGUF V3 (latest) print_info: file type = STQ1_0 - 1.31 bpw ternary print_info: file size = 435.61 MiB (2.04 BPW) load: 0 unused tokens load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect load: printing all EOG tokens: load: - 120020 ('<|hy_place▁holder▁no▁2|>') load: special tokens cache size = 818 load: token to piece cache size = 0.8089 MB print_info: arch = hunyuan-dense print_info: vocab_only = 0 print_info: no_alloc = 0 print_info: n_ctx_train = 262144 print_info: n_embd = 2048 print_info: n_embd_inp = 2048 print_info: n_layer = 32 print_info: n_head = 16 print_info: n_head_kv = 4 print_info: n_rot = 128 print_info: n_swa = 0 print_info: is_swa_any = 0 print_info: n_embd_head_k = 128 print_info: n_embd_head_v = 128 print_info: n_gqa = 4 print_info: n_embd_k_gqa = 512 print_info: n_embd_v_gqa = 512 print_info: f_norm_eps = 0.0e+00 print_info: f_norm_rms_eps = 1.0e-05 print_info: f_clamp_kqv = 0.0e+00 print_info: f_max_alibi_bias = 0.0e+00 print_info: f_logit_scale = 0.0e+00 print_info: f_attn_scale = 0.0e+00 print_info: f_attn_value_scale = 0.0000 print_info: n_ff = 6144 print_info: n_expert = 0 print_info: n_expert_used = 0 print_info: n_expert_groups = 0 print_info: n_group_used = 0 print_info: causal attn = 1 print_info: pooling type = -1 print_info: rope type = 2 print_info: rope scaling = none print_info: freq_base_train = 11158840.0 print_info: freq_scale_train = 1 print_info: n_ctx_orig_yarn = 262144 print_info: rope_yarn_log_mul = 0.0000 print_info: rope_finetuned = unknown print_info: model type = 1.8B print_info: model params = 1.79 B print_info: general.name = Hy MT1.5 1.8B 1.25bit print_info: vocab type = BPE print_info: n_vocab = 120818 print_info: n_merges = 119758 print_info: BOS token = 120000 '<|hy_begin▁of▁sentence|>' print_info: EOS token = 120020 '<|hy_place▁holder▁no▁2|>' print_info: SEP token = 120007 '<|hy_Assistant|>' print_info: PAD token = 120002 '<|hy_▁pad▁|>' print_info: LF token = 185 'Ċ' print_info: EOG token = 120020 '<|hy_place▁holder▁no▁2|>' print_info: max token length = 1024 load_tensors: loading model tensors, this can take a while... (mmap = true, direct_io = false) load_tensors: CPU_Mapped model buffer size = 435.61 MiB ......................................................... common_init_result: added <|hy_place▁holder▁no▁2|> logit bias = -inf llama_context: constructing llama_context llama_context: n_seq_max = 1 llama_context: n_ctx = 219904 llama_context: n_ctx_seq = 219904 llama_context: n_batch = 2048 llama_context: n_ubatch = 512 llama_context: causal_attn = 1 llama_context: flash_attn = auto llama_context: kv_unified = false llama_context: freq_base = 11158840.0 llama_context: freq_scale = 1 llama_context: n_ctx_seq (219904) < n_ctx_train (262144) -- the full capacity of the model will not be utilized llama_context: CPU output buffer size = 0.46 MiB llama_kv_cache: CPU KV buffer size = 13744.00 MiB Killed

still error

Looks like it's being killed due to OOM. The default context (262144) makes the KV cache ~16 GiB, which may exceed host memory.

Could you try again with -c 4096 ? That brings the KV cache down to ~256 MiB, plenty for a short translation:

./build/bin/llama-completion --model ../model_zoo/Hy-MT1.5-1.8B-1.25bit-GGUF/Hy-MT1.5-1.8B-1.25bit.gguf \
-p "Translate the following segment into Chinese, without additional explanation:Hello" \
--jinja -ngl 0 -c 4096 -n 64 -st

@tritueviet

Copy link
Copy Markdown

./build/bin/llama-completion --model ../model_zoo/Hy-MT1.5-1.8B-1.25bit-GGUF/Hy-MT1.5-1.8B-1.25bit.gguf -p "Translate the following segment into Chinese, without additional explanation:Hello" --jinja -ngl 0 -n 64 -st warning: no usable GPU found, --gpu-layers option will be ignored warning: one possible reason is that llama.cpp was compiled without GPU support warning: consult docs/build.md for compilation instructions main: llama backend init main: load the model and apply lora adapter, if any common_init_result: fitting params to device memory, for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on common_params_fit_impl: getting device memory data for initial parameters: common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | common_memory_breakdown_print: | - Host | 17607 = 435 + 16384 + 788 | common_params_fit_impl: projected to use 17607 MiB of host memory vs. 15908 MiB of total host memory common_params_fit_impl: cannot meet free memory target of 1024 MiB, need to reduce device memory by 2723 MiB common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | common_memory_breakdown_print: | - Host | 939 = 435 + 256 + 247 | common_params_fit_impl: context size reduced from 262144 to 219904 -> need 2728 MiB less memory in total common_params_fit_impl: entire model can be fit by reducing context common_fit_params: successfully fit params to free device memory common_fit_params: fitting params to free memory took 0.55 seconds llama_model_loader: loaded meta data with 40 key-value pairs and 354 tensors from ../model_zoo/Hy-MT1.5-1.8B-1.25bit-GGUF/Hy-MT1.5-1.8B-1.25bit.gguf (version GGUF V3 (latest)) llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output. llama_model_loader: - kv 0: general.architecture str = hunyuan-dense llama_model_loader: - kv 1: general.type str = model llama_model_loader: - kv 2: general.sampling.top_k i32 = 20 llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.800000 llama_model_loader: - kv 4: general.sampling.temp f32 = 0.700000 llama_model_loader: - kv 5: general.name str = Hy MT1.5 1.8B 1.25bit llama_model_loader: - kv 6: general.finetune str = 1.25bit llama_model_loader: - kv 7: general.basename str = Hy-MT1.5 llama_model_loader: - kv 8: general.size_label str = 1.8B llama_model_loader: - kv 9: general.base_model.count u32 = 1 llama_model_loader: - kv 10: general.base_model.0.name str = HY MT1.5 1.8B llama_model_loader: - kv 11: general.base_model.0.organization str = Tencent llama_model_loader: - kv 12: general.base_model.0.repo_url str = https://huggingface.co/tencent/HY-MT1... llama_model_loader: - kv 13: general.tags arr[str,5] = ["translation", "hy-mt", "quant", "1.... llama_model_loader: - kv 14: general.languages arr[str,1] = ["multilingual"] llama_model_loader: - kv 15: hunyuan-dense.block_count u32 = 32 llama_model_loader: - kv 16: hunyuan-dense.context_length u32 = 262144 llama_model_loader: - kv 17: hunyuan-dense.embedding_length u32 = 2048 llama_model_loader: - kv 18: hunyuan-dense.feed_forward_length u32 = 6144 llama_model_loader: - kv 19: hunyuan-dense.attention.head_count u32 = 16 llama_model_loader: - kv 20: hunyuan-dense.attention.head_count_kv u32 = 4 llama_model_loader: - kv 21: hunyuan-dense.rope.freq_base f32 = 11158840.000000 llama_model_loader: - kv 22: hunyuan-dense.attention.layer_norm_rms_epsilon f32 = 0.000010 llama_model_loader: - kv 23: hunyuan-dense.attention.key_length u32 = 128 llama_model_loader: - kv 24: hunyuan-dense.attention.value_length u32 = 128 llama_model_loader: - kv 25: hunyuan-dense.rope.scaling.type str = none llama_model_loader: - kv 26: hunyuan-dense.rope.scaling.factor f32 = 1.000000 llama_model_loader: - kv 27: hunyuan-dense.rope.scaling.original_context_length u32 = 262144 llama_model_loader: - kv 28: tokenizer.ggml.model str = gpt2 llama_model_loader: - kv 29: tokenizer.ggml.pre str = hunyuan-dense llama_model_loader: - kv 30: tokenizer.ggml.tokens arr[str,120818] = ["!", """, "#", "$", "%", "&", "'", ... llama_model_loader: - kv 31: tokenizer.ggml.token_type arr[i32,120818] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ... llama_model_loader: - kv 32: tokenizer.ggml.merges arr[str,119758] = ["Ġ Ġ", "Ġ t", "Ġ a", "i n", "h e... llama_model_loader: - kv 33: tokenizer.ggml.bos_token_id u32 = 120000 llama_model_loader: - kv 34: tokenizer.ggml.eos_token_id u32 = 120020 llama_model_loader: - kv 35: tokenizer.ggml.padding_token_id u32 = 120002 llama_model_loader: - kv 36: tokenizer.ggml.seperator_token_id u32 = 120007 llama_model_loader: - kv 37: tokenizer.chat_template str = {% if messages[0]['role'] == 'system'... llama_model_loader: - kv 38: general.quantization_version u32 = 2 llama_model_loader: - kv 39: general.file_type u32 = 41 llama_model_loader: - type f32: 129 tensors llama_model_loader: - type q6_K: 1 tensors llama_model_loader: - type stq1_0: 224 tensors print_info: file format = GGUF V3 (latest) print_info: file type = STQ1_0 - 1.31 bpw ternary print_info: file size = 435.61 MiB (2.04 BPW) load: 0 unused tokens load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect load: printing all EOG tokens: load: - 120020 ('<|hy_place▁holder▁no▁2|>') load: special tokens cache size = 818 load: token to piece cache size = 0.8089 MB print_info: arch = hunyuan-dense print_info: vocab_only = 0 print_info: no_alloc = 0 print_info: n_ctx_train = 262144 print_info: n_embd = 2048 print_info: n_embd_inp = 2048 print_info: n_layer = 32 print_info: n_head = 16 print_info: n_head_kv = 4 print_info: n_rot = 128 print_info: n_swa = 0 print_info: is_swa_any = 0 print_info: n_embd_head_k = 128 print_info: n_embd_head_v = 128 print_info: n_gqa = 4 print_info: n_embd_k_gqa = 512 print_info: n_embd_v_gqa = 512 print_info: f_norm_eps = 0.0e+00 print_info: f_norm_rms_eps = 1.0e-05 print_info: f_clamp_kqv = 0.0e+00 print_info: f_max_alibi_bias = 0.0e+00 print_info: f_logit_scale = 0.0e+00 print_info: f_attn_scale = 0.0e+00 print_info: f_attn_value_scale = 0.0000 print_info: n_ff = 6144 print_info: n_expert = 0 print_info: n_expert_used = 0 print_info: n_expert_groups = 0 print_info: n_group_used = 0 print_info: causal attn = 1 print_info: pooling type = -1 print_info: rope type = 2 print_info: rope scaling = none print_info: freq_base_train = 11158840.0 print_info: freq_scale_train = 1 print_info: n_ctx_orig_yarn = 262144 print_info: rope_yarn_log_mul = 0.0000 print_info: rope_finetuned = unknown print_info: model type = 1.8B print_info: model params = 1.79 B print_info: general.name = Hy MT1.5 1.8B 1.25bit print_info: vocab type = BPE print_info: n_vocab = 120818 print_info: n_merges = 119758 print_info: BOS token = 120000 '<|hy_begin▁of▁sentence|>' print_info: EOS token = 120020 '<|hy_place▁holder▁no▁2|>' print_info: SEP token = 120007 '<|hy_Assistant|>' print_info: PAD token = 120002 '<|hy_▁pad▁|>' print_info: LF token = 185 'Ċ' print_info: EOG token = 120020 '<|hy_place▁holder▁no▁2|>' print_info: max token length = 1024 load_tensors: loading model tensors, this can take a while... (mmap = true, direct_io = false) load_tensors: CPU_Mapped model buffer size = 435.61 MiB ......................................................... common_init_result: added <|hy_place▁holder▁no▁2|> logit bias = -inf llama_context: constructing llama_context llama_context: n_seq_max = 1 llama_context: n_ctx = 219904 llama_context: n_ctx_seq = 219904 llama_context: n_batch = 2048 llama_context: n_ubatch = 512 llama_context: causal_attn = 1 llama_context: flash_attn = auto llama_context: kv_unified = false llama_context: freq_base = 11158840.0 llama_context: freq_scale = 1 llama_context: n_ctx_seq (219904) < n_ctx_train (262144) -- the full capacity of the model will not be utilized llama_context: CPU output buffer size = 0.46 MiB llama_kv_cache: CPU KV buffer size = 13744.00 MiB Killed
still error

Looks like it's being killed due to OOM. The default context (262144) makes the KV cache ~16 GiB, which may exceed host memory.

Could you try again with -c 4096 ? That brings the KV cache down to ~256 MiB, plenty for a short translation:

./build/bin/llama-completion --model ../model_zoo/Hy-MT1.5-1.8B-1.25bit-GGUF/Hy-MT1.5-1.8B-1.25bit.gguf \
-p "Translate the following segment into Chinese, without additional explanation:Hello" \
--jinja -ngl 0 -c 4096 -n 64 -st

./build/bin/llama-completion --model model_zoo/Hy-MT1.5-1.8B-STQ1_0.gguf -p "Translate the following segment into Vietnamese, without additional explanation. Hello " --jinja -ngl 0 -c 4096 -n 64 -st
warning: no usable GPU found, --gpu-layers option will be ignored
warning: one possible reason is that llama.cpp was compiled without GPU support
warning: consult docs/build.md for compilation instructions
main: llama backend init
main: load the model and apply lora adapter, if any
common_init_result: fitting params to device memory, for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on
common_params_fit_impl: getting device memory data for initial parameters:
common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted |
common_memory_breakdown_print: | - Host | 939 = 435 + 256 + 247 |
common_params_fit_impl: projected to use 939 MiB of host memory vs. 15625 MiB of total host memory
common_params_fit_impl: will leave 14686 >= 1024 MiB of system memory, no changes needed
common_fit_params: successfully fit params to free device memory
common_fit_params: fitting params to free memory took 0.21 seconds
llama_model_loader: loaded meta data with 40 key-value pairs and 354 tensors from model_zoo/Hy-MT1.5-1.8B-STQ1_0.gguf (version GGUF V3 (latest))
llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output.
llama_model_loader: - kv 0: general.architecture str = hunyuan-dense
llama_model_loader: - kv 1: general.type str = model
llama_model_loader: - kv 2: general.sampling.top_k i32 = 20
llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.800000
llama_model_loader: - kv 4: general.sampling.temp f32 = 0.700000
llama_model_loader: - kv 5: general.name str = Hy MT1.5 1.8B 1.25bit
llama_model_loader: - kv 6: general.finetune str = 1.25bit
llama_model_loader: - kv 7: general.basename str = Hy-MT1.5
llama_model_loader: - kv 8: general.size_label str = 1.8B
llama_model_loader: - kv 9: general.base_model.count u32 = 1
llama_model_loader: - kv 10: general.base_model.0.name str = HY MT1.5 1.8B
llama_model_loader: - kv 11: general.base_model.0.organization str = Tencent
llama_model_loader: - kv 12: general.base_model.0.repo_url str = https://huggingface.co/tencent/HY-MT1...
llama_model_loader: - kv 13: general.tags arr[str,5] = ["translation", "hy-mt", "quant", "1....
llama_model_loader: - kv 14: general.languages arr[str,1] = ["multilingual"]
llama_model_loader: - kv 15: hunyuan-dense.block_count u32 = 32
llama_model_loader: - kv 16: hunyuan-dense.context_length u32 = 262144
llama_model_loader: - kv 17: hunyuan-dense.embedding_length u32 = 2048
llama_model_loader: - kv 18: hunyuan-dense.feed_forward_length u32 = 6144
llama_model_loader: - kv 19: hunyuan-dense.attention.head_count u32 = 16
llama_model_loader: - kv 20: hunyuan-dense.attention.head_count_kv u32 = 4
llama_model_loader: - kv 21: hunyuan-dense.rope.freq_base f32 = 11158840.000000
llama_model_loader: - kv 22: hunyuan-dense.attention.layer_norm_rms_epsilon f32 = 0.000010
llama_model_loader: - kv 23: hunyuan-dense.attention.key_length u32 = 128
llama_model_loader: - kv 24: hunyuan-dense.attention.value_length u32 = 128
llama_model_loader: - kv 25: hunyuan-dense.rope.scaling.type str = none
llama_model_loader: - kv 26: hunyuan-dense.rope.scaling.factor f32 = 1.000000
llama_model_loader: - kv 27: hunyuan-dense.rope.scaling.original_context_length u32 = 262144
llama_model_loader: - kv 28: tokenizer.ggml.model str = gpt2
llama_model_loader: - kv 29: tokenizer.ggml.pre str = hunyuan-dense
llama_model_loader: - kv 30: tokenizer.ggml.tokens arr[str,120818] = ["!", """, "#", "$", "%", "&", "'", ...
llama_model_loader: - kv 31: tokenizer.ggml.token_type arr[i32,120818] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ...
llama_model_loader: - kv 32: tokenizer.ggml.merges arr[str,119758] = ["Ġ Ġ", "Ġ t", "Ġ a", "i n", "h e...
llama_model_loader: - kv 33: tokenizer.ggml.bos_token_id u32 = 120000
llama_model_loader: - kv 34: tokenizer.ggml.eos_token_id u32 = 3
llama_model_loader: - kv 35: tokenizer.ggml.padding_token_id u32 = 120002
llama_model_loader: - kv 36: tokenizer.ggml.seperator_token_id u32 = 120007
llama_model_loader: - kv 37: tokenizer.chat_template str = {% if messages[0]['role'] == 'system'...
llama_model_loader: - kv 38: general.quantization_version u32 = 2
llama_model_loader: - kv 39: general.file_type u32 = 41
llama_model_loader: - type f32: 129 tensors
llama_model_loader: - type q6_K: 1 tensors
llama_model_loader: - type stq1_0: 224 tensors
print_info: file format = GGUF V3 (latest)
print_info: file type = STQ1_0 - 1.31 bpw ternary
print_info: file size = 435.61 MiB (2.04 BPW)
load: 0 unused tokens
load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect
load: printing all EOG tokens:
load: - 3 ('$')
load: special tokens cache size = 818
load: token to piece cache size = 0.8089 MB
print_info: arch = hunyuan-dense
print_info: vocab_only = 0
print_info: no_alloc = 0
print_info: n_ctx_train = 262144
print_info: n_embd = 2048
print_info: n_embd_inp = 2048
print_info: n_layer = 32
print_info: n_head = 16
print_info: n_head_kv = 4
print_info: n_rot = 128
print_info: n_swa = 0
print_info: is_swa_any = 0
print_info: n_embd_head_k = 128
print_info: n_embd_head_v = 128
print_info: n_gqa = 4
print_info: n_embd_k_gqa = 512
print_info: n_embd_v_gqa = 512
print_info: f_norm_eps = 0.0e+00
print_info: f_norm_rms_eps = 1.0e-05
print_info: f_clamp_kqv = 0.0e+00
print_info: f_max_alibi_bias = 0.0e+00
print_info: f_logit_scale = 0.0e+00
print_info: f_attn_scale = 0.0e+00
print_info: f_attn_value_scale = 0.0000
print_info: n_ff = 6144
print_info: n_expert = 0
print_info: n_expert_used = 0
print_info: n_expert_groups = 0
print_info: n_group_used = 0
print_info: causal attn = 1
print_info: pooling type = -1
print_info: rope type = 2
print_info: rope scaling = none
print_info: freq_base_train = 11158840.0
print_info: freq_scale_train = 1
print_info: n_ctx_orig_yarn = 262144
print_info: rope_yarn_log_mul = 0.0000
print_info: rope_finetuned = unknown
print_info: model type = 1.8B
print_info: model params = 1.79 B
print_info: general.name = Hy MT1.5 1.8B 1.25bit
print_info: vocab type = BPE
print_info: n_vocab = 120818
print_info: n_merges = 119758
print_info: BOS token = 120000 '<|hy_begin▁of▁sentence|>'
print_info: EOS token = 3 '$'
print_info: SEP token = 120007 '<|hy_Assistant|>'
print_info: PAD token = 120002 '<|hy_▁pad▁|>'
print_info: LF token = 185 'Ċ'
print_info: EOG token = 3 '$'
print_info: max token length = 1024
load_tensors: loading model tensors, this can take a while... (mmap = true, direct_io = false)
load_tensors: CPU_Mapped model buffer size = 435.61 MiB
.........................................................
common_init_result: added $ logit bias = -inf
llama_context: constructing llama_context
llama_context: n_seq_max = 1
llama_context: n_ctx = 4096
llama_context: n_ctx_seq = 4096
llama_context: n_batch = 2048
llama_context: n_ubatch = 512
llama_context: causal_attn = 1
llama_context: flash_attn = auto
llama_context: kv_unified = false
llama_context: freq_base = 11158840.0
llama_context: freq_scale = 1
llama_context: n_ctx_seq (4096) < n_ctx_train (262144) -- the full capacity of the model will not be utilized
llama_context: CPU output buffer size = 0.46 MiB
llama_kv_cache: CPU KV buffer size = 256.00 MiB
llama_kv_cache: size = 256.00 MiB ( 4096 cells, 32 layers, 1/1 seqs), K (f16): 128.00 MiB, V (f16): 128.00 MiB
llama_kv_cache: attn_rot_k = 0, n_embd_head_k_all = 128
llama_kv_cache: attn_rot_v = 0, n_embd_head_k_all = 128
sched_reserve: reserving ...
sched_reserve: Flash Attention was auto, set to enabled
sched_reserve: resolving fused Gated Delta Net support:
sched_reserve: fused Gated Delta Net (autoregressive) enabled
sched_reserve: fused Gated Delta Net (chunked) enabled
sched_reserve: CPU compute buffer size = 247.97 MiB
sched_reserve: graph nodes = 1127
sched_reserve: graph splits = 1
sched_reserve: reserve took 2.11 ms, sched copies = 1
common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable)
main: llama threadpool init, n_threads = 4
main: chat template is available, enabling conversation mode (disable it with -no-cnv)
*** User-specified prompt will pre-start conversation, did you mean to set --system-prompt (-sys) instead?
main: chat template example:
<|hy_begin▁of▁sentence|>You are a helpful assistant<|hy_place▁holder▁no▁3|><|hy_User|>Hello<|hy_Assistant|>Hi there<|hy_place▁holder▁no▁2|><|hy_User|>How are you?<|hy_Assistant|>

system_info: n_threads = 4 (n_threads_batch = 4) / 16 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |

sampler seed: 3567722247
sampler params:
repeat_last_n = 64, repeat_penalty = 1.000, frequency_penalty = 0.000, presence_penalty = 0.000
dry_multiplier = 0.000, dry_base = 1.750, dry_allowed_length = 2, dry_penalty_last_n = -1
top_k = 20, top_p = 0.800, min_p = 0.050, xtc_probability = 0.000, xtc_threshold = 0.100, typical_p = 1.000, top_n_sigma = -1.000, temp = 0.700
mirostat = 0, mirostat_lr = 0.100, mirostat_ent = 5.000, adaptive_target = -1.000, adaptive_decay = 0.900
sampler chain: logits -> ?penalties -> ?dry -> ?top-n-sigma -> top-k -> ?typical -> top-p -> min-p -> ?xtc -> temp-ext -> dist
generate: n_ctx = 4096, n_batch = 2048, n_predict = 64, n_keep = 0

Translate the following segment into Vietnamese, without additional explanation. Hello Xin chào.َايْتُمْ بَعْضُنُمْ بَعْلَيٰ ٌبَعْلَيٰ ٌبَعْلَيٰ… الْمُنْ

common_perf_print: sampling time = 13.33 ms
common_perf_print: samplers time = 4.08 ms / 80 tokens
common_perf_print: load time = 830.89 ms
common_perf_print: prompt eval time = 4459.01 ms / 16 tokens ( 278.69 ms per token, 3.59 tokens per second)
common_perf_print: eval time = 18632.27 ms / 63 runs ( 295.75 ms per token, 3.38 tokens per second)
common_perf_print: total time = 23108.99 ms / 79 tokens
common_perf_print: unaccounted time = 4.38 ms / 0.0 % (total - sampling - prompt eval - eval) / (total)
common_perf_print: graphs reused = 62
common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted |
common_memory_breakdown_print: | - Host | 939 = 435 + 256 + 247 |

performance and result not good. it only run with arm??

@sjl623

sjl623 commented May 12, 2026 •

Copy link
Copy Markdown
Author

./build/bin/llama-completion --model model_zoo/Hy-MT1.5-1.8B-STQ1_0.gguf -p "Translate the following segment into Vietnamese, without additional explanation. Hello " --jinja -ngl 0 -c 4096 -n 64 -st warning: no usable GPU found, --gpu-layers option will be ignored warning: one possible reason is that llama.cpp was compiled without GPU support warning: consult docs/build.md for compilation instructions main: llama backend init main: load the model and apply lora adapter, if any common_init_result: fitting params to device memory, for bugs during this step try to reproduce them with -fit off, or provide --verbose logs if the bug only occurs with -fit on common_params_fit_impl: getting device memory data for initial parameters: common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | common_memory_breakdown_print: | - Host | 939 = 435 + 256 + 247 | common_params_fit_impl: projected to use 939 MiB of host memory vs. 15625 MiB of total host memory common_params_fit_impl: will leave 14686 >= 1024 MiB of system memory, no changes needed common_fit_params: successfully fit params to free device memory common_fit_params: fitting params to free memory took 0.21 seconds llama_model_loader: loaded meta data with 40 key-value pairs and 354 tensors from model_zoo/Hy-MT1.5-1.8B-STQ1_0.gguf (version GGUF V3 (latest)) llama_model_loader: Dumping metadata keys/values. Note: KV overrides do not apply in this output. llama_model_loader: - kv 0: general.architecture str = hunyuan-dense llama_model_loader: - kv 1: general.type str = model llama_model_loader: - kv 2: general.sampling.top_k i32 = 20 llama_model_loader: - kv 3: general.sampling.top_p f32 = 0.800000 llama_model_loader: - kv 4: general.sampling.temp f32 = 0.700000 llama_model_loader: - kv 5: general.name str = Hy MT1.5 1.8B 1.25bit llama_model_loader: - kv 6: general.finetune str = 1.25bit llama_model_loader: - kv 7: general.basename str = Hy-MT1.5 llama_model_loader: - kv 8: general.size_label str = 1.8B llama_model_loader: - kv 9: general.base_model.count u32 = 1 llama_model_loader: - kv 10: general.base_model.0.name str = HY MT1.5 1.8B llama_model_loader: - kv 11: general.base_model.0.organization str = Tencent llama_model_loader: - kv 12: general.base_model.0.repo_url str = https://huggingface.co/tencent/HY-MT1... llama_model_loader: - kv 13: general.tags arr[str,5] = ["translation", "hy-mt", "quant", "1.... llama_model_loader: - kv 14: general.languages arr[str,1] = ["multilingual"] llama_model_loader: - kv 15: hunyuan-dense.block_count u32 = 32 llama_model_loader: - kv 16: hunyuan-dense.context_length u32 = 262144 llama_model_loader: - kv 17: hunyuan-dense.embedding_length u32 = 2048 llama_model_loader: - kv 18: hunyuan-dense.feed_forward_length u32 = 6144 llama_model_loader: - kv 19: hunyuan-dense.attention.head_count u32 = 16 llama_model_loader: - kv 20: hunyuan-dense.attention.head_count_kv u32 = 4 llama_model_loader: - kv 21: hunyuan-dense.rope.freq_base f32 = 11158840.000000 llama_model_loader: - kv 22: hunyuan-dense.attention.layer_norm_rms_epsilon f32 = 0.000010 llama_model_loader: - kv 23: hunyuan-dense.attention.key_length u32 = 128 llama_model_loader: - kv 24: hunyuan-dense.attention.value_length u32 = 128 llama_model_loader: - kv 25: hunyuan-dense.rope.scaling.type str = none llama_model_loader: - kv 26: hunyuan-dense.rope.scaling.factor f32 = 1.000000 llama_model_loader: - kv 27: hunyuan-dense.rope.scaling.original_context_length u32 = 262144 llama_model_loader: - kv 28: tokenizer.ggml.model str = gpt2 llama_model_loader: - kv 29: tokenizer.ggml.pre str = hunyuan-dense llama_model_loader: - kv 30: tokenizer.ggml.tokens arr[str,120818] = ["!", """, "#", "$", "%", "&", "'", ... llama_model_loader: - kv 31: tokenizer.ggml.token_type arr[i32,120818] = [1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, ... llama_model_loader: - kv 32: tokenizer.ggml.merges arr[str,119758] = ["Ġ Ġ", "Ġ t", "Ġ a", "i n", "h e... llama_model_loader: - kv 33: tokenizer.ggml.bos_token_id u32 = 120000 llama_model_loader: - kv 34: tokenizer.ggml.eos_token_id u32 = 3 llama_model_loader: - kv 35: tokenizer.ggml.padding_token_id u32 = 120002 llama_model_loader: - kv 36: tokenizer.ggml.seperator_token_id u32 = 120007 llama_model_loader: - kv 37: tokenizer.chat_template str = {% if messages[0]['role'] == 'system'... llama_model_loader: - kv 38: general.quantization_version u32 = 2 llama_model_loader: - kv 39: general.file_type u32 = 41 llama_model_loader: - type f32: 129 tensors llama_model_loader: - type q6_K: 1 tensors llama_model_loader: - type stq1_0: 224 tensors print_info: file format = GGUF V3 (latest) print_info: file type = STQ1_0 - 1.31 bpw ternary print_info: file size = 435.61 MiB (2.04 BPW) load: 0 unused tokens load: special_eos_id is not in special_eog_ids - the tokenizer config may be incorrect load: printing all EOG tokens: load: - 3 ('$') load: special tokens cache size = 818 load: token to piece cache size = 0.8089 MB print_info: arch = hunyuan-dense print_info: vocab_only = 0 print_info: no_alloc = 0 print_info: n_ctx_train = 262144 print_info: n_embd = 2048 print_info: n_embd_inp = 2048 print_info: n_layer = 32 print_info: n_head = 16 print_info: n_head_kv = 4 print_info: n_rot = 128 print_info: n_swa = 0 print_info: is_swa_any = 0 print_info: n_embd_head_k = 128 print_info: n_embd_head_v = 128 print_info: n_gqa = 4 print_info: n_embd_k_gqa = 512 print_info: n_embd_v_gqa = 512 print_info: f_norm_eps = 0.0e+00 print_info: f_norm_rms_eps = 1.0e-05 print_info: f_clamp_kqv = 0.0e+00 print_info: f_max_alibi_bias = 0.0e+00 print_info: f_logit_scale = 0.0e+00 print_info: f_attn_scale = 0.0e+00 print_info: f_attn_value_scale = 0.0000 print_info: n_ff = 6144 print_info: n_expert = 0 print_info: n_expert_used = 0 print_info: n_expert_groups = 0 print_info: n_group_used = 0 print_info: causal attn = 1 print_info: pooling type = -1 print_info: rope type = 2 print_info: rope scaling = none print_info: freq_base_train = 11158840.0 print_info: freq_scale_train = 1 print_info: n_ctx_orig_yarn = 262144 print_info: rope_yarn_log_mul = 0.0000 print_info: rope_finetuned = unknown print_info: model type = 1.8B print_info: model params = 1.79 B print_info: general.name = Hy MT1.5 1.8B 1.25bit print_info: vocab type = BPE print_info: n_vocab = 120818 print_info: n_merges = 119758 print_info: BOS token = 120000 '<|hy_begin▁of▁sentence|>' print_info: EOS token = 3 '$' print_info: SEP token = 120007 '<|hy_Assistant|>' print_info: PAD token = 120002 '<|hy_▁pad▁|>' print_info: LF token = 185 'Ċ' print_info: EOG token = 3 '$' print_info: max token length = 1024 load_tensors: loading model tensors, this can take a while... (mmap = true, direct_io = false) load_tensors: CPU_Mapped model buffer size = 435.61 MiB ......................................................... common_init_result: added $ logit bias = -inf llama_context: constructing llama_context llama_context: n_seq_max = 1 llama_context: n_ctx = 4096 llama_context: n_ctx_seq = 4096 llama_context: n_batch = 2048 llama_context: n_ubatch = 512 llama_context: causal_attn = 1 llama_context: flash_attn = auto llama_context: kv_unified = false llama_context: freq_base = 11158840.0 llama_context: freq_scale = 1 llama_context: n_ctx_seq (4096) < n_ctx_train (262144) -- the full capacity of the model will not be utilized llama_context: CPU output buffer size = 0.46 MiB llama_kv_cache: CPU KV buffer size = 256.00 MiB llama_kv_cache: size = 256.00 MiB ( 4096 cells, 32 layers, 1/1 seqs), K (f16): 128.00 MiB, V (f16): 128.00 MiB llama_kv_cache: attn_rot_k = 0, n_embd_head_k_all = 128 llama_kv_cache: attn_rot_v = 0, n_embd_head_k_all = 128 sched_reserve: reserving ... sched_reserve: Flash Attention was auto, set to enabled sched_reserve: resolving fused Gated Delta Net support: sched_reserve: fused Gated Delta Net (autoregressive) enabled sched_reserve: fused Gated Delta Net (chunked) enabled sched_reserve: CPU compute buffer size = 247.97 MiB sched_reserve: graph nodes = 1127 sched_reserve: graph splits = 1 sched_reserve: reserve took 2.11 ms, sched copies = 1 common_init_from_params: warming up the model with an empty run - please wait ... (--no-warmup to disable) main: llama threadpool init, n_threads = 4 main: chat template is available, enabling conversation mode (disable it with -no-cnv) *** User-specified prompt will pre-start conversation, did you mean to set --system-prompt (-sys) instead? main: chat template example: <|hy_begin▁of▁sentence|>You are a helpful assistant<|hy_place▁holder▁no▁3|><|hy_User|>Hello<|hy_Assistant|>Hi there<|hy_place▁holder▁no▁2|><|hy_User|>How are you?<|hy_Assistant|>

system_info: n_threads = 4 (n_threads_batch = 4) / 16 | CPU : SSE3 = 1 | SSSE3 = 1 | AVX = 1 | AVX_VNNI = 1 | AVX2 = 1 | F16C = 1 | FMA = 1 | BMI2 = 1 | LLAMAFILE = 1 | OPENMP = 1 | REPACK = 1 |

sampler seed: 3567722247 sampler params: repeat_last_n = 64, repeat_penalty = 1.000, frequency_penalty = 0.000, presence_penalty = 0.000 dry_multiplier = 0.000, dry_base = 1.750, dry_allowed_length = 2, dry_penalty_last_n = -1 top_k = 20, top_p = 0.800, min_p = 0.050, xtc_probability = 0.000, xtc_threshold = 0.100, typical_p = 1.000, top_n_sigma = -1.000, temp = 0.700 mirostat = 0, mirostat_lr = 0.100, mirostat_ent = 5.000, adaptive_target = -1.000, adaptive_decay = 0.900 sampler chain: logits -> ?penalties -> ?dry -> ?top-n-sigma -> top-k -> ?typical -> top-p -> min-p -> ?xtc -> temp-ext -> dist generate: n_ctx = 4096, n_batch = 2048, n_predict = 64, n_keep = 0

Translate the following segment into Vietnamese, without additional explanation. Hello Xin chào.َايْتُمْ بَعْضُنُمْ بَعْلَيٰ ٌبَعْلَيٰ ٌبَعْلَيٰ… الْمُنْ

common_perf_print: sampling time = 13.33 ms common_perf_print: samplers time = 4.08 ms / 80 tokens common_perf_print: load time = 830.89 ms common_perf_print: prompt eval time = 4459.01 ms / 16 tokens ( 278.69 ms per token, 3.59 tokens per second) common_perf_print: eval time = 18632.27 ms / 63 runs ( 295.75 ms per token, 3.38 tokens per second) common_perf_print: total time = 23108.99 ms / 79 tokens common_perf_print: unaccounted time = 4.38 ms / 0.0 % (total - sampling - prompt eval - eval) / (total) common_perf_print: graphs reused = 62 common_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | common_memory_breakdown_print: | - Host | 939 = 435 + 256 + 247 |

performance and result not good. it only run with arm??

Yes, currently the acceleration only supports ARM chips. You can try it on an M-series MacBook, or download our APK to experience it directly.

@sjl623 sjl623 closed this May 12, 2026
@sjl623 sjl623 reopened this May 12, 2026
@tritueviet

Copy link
Copy Markdown

approved.

justinchuby added a commit to onnxruntime/mobius that referenced this pull request May 20, 2026
## Summary

Adds two paths through the GGUF loader for GGML type 41 so mobius can
ingest both standard and Tencent's custom variants of Q1_0.

### 1. Mainline Q1_0 — 1-bit binary

Layout (18 bytes / 128 elements): `[fp16 d][16B packed 1-bit signs]`.
Dequant: `bit ? +d : -d`. Used by
[prism-ml/Bonsai](https://huggingface.co/collections/prism-ml/bonsai)
checkpoints.

Mapped to `MatMulNBits` `bits=2`, `zero_point=1`, codes ∈ {0, 2},
scale=d.

### 2. Tencent custom Q1_0 — 2-bit SEQ

Used by
[`AngelSlim/Hy-MT1.5-1.8B-2bit-GGUF`](https://huggingface.co/AngelSlim/Hy-MT1.5-1.8B-2bit-GGUF).
Per-row layout: `[fp16 stored_scale][2-bit codes packed LSB-first]`,
native block size 512, codebook `{−3, −1, +1, +3} · stored_scale`.

Mainline llama.cpp **refuses to load** these files because every tensor
after the first Q1_0 entry lands at the wrong offset. We bypass the size
calculation by reading each tensor from its explicit GGUF offset.

Two `MatMulNBits` representations, selected by a flag:

| | default | `MOBIUS_TENCENT_Q1_0_USE_NATIVE_2BIT=1` |
|---|---|---|
| ORT bits | 4 (inflated `2c ∈ {0,2,4,6}`, integer zp=3) | 2 (codes
pass-through, float zp=1.5) |
| on-disk weight bpw | 4 (2× the source) | 2 (matches source) |
| CPU decode throughput | **~32 tok/s** | **~0.24 tok/s** |
| ORT version | any | **≥ 1.27**
([#28354](microsoft/onnxruntime#28354)) |

Both representations produce **bit-identical dequantized weights**
matching HF safetensors to bf16 rounding (`max_abs ≈ 2.5e-3` vs HF
reference matmul). The 130× decode speed difference is purely the ORT
CPU kernel: the bits=4 path is the mature MLAS fused path, the bits=2 +
float-zp path is the kernel's own self-described "naive implementation,
need to be optimized" scalar fallback. Tracked in
[microsoft/onnxruntime#28552](microsoft/onnxruntime#28552);
once that lands the default can flip.

**Block-size workaround:** ORT silently returns zeros for `block_size >
256` ([#28551](microsoft/onnxruntime#28551)).
Both representations therefore expose `block_size = 128` by replicating
each native 512-element scale across 4 sub-blocks.

## Other infrastructure changes

- New flag [`tencent_q1_0_use_native_2bit`](src/mobius/_flags.py)
(default `False`).
- `QuantizedLinear` accepts `bits ∈ {2, 4, 8}` and a new
`zero_point_dtype` argument. `UINT8` (default) keeps the bit-packed
integer form; float dtypes produce one un-packed value per block.
- `QuantizationConfig` gains `float_zero_point`; `TextModel` passes
`config.dtype` as the zp dtype when it is set.
- GGUF → `ArchitectureConfig` derivation now sets `rope_type` from
`<arch>.rope.scaling.type` (defaulting to `"default"` when
`rope.freq_base` is present). Without this fix, models with
`rope.scaling.type = "none"` were built with no RoPE at all.
- For `hunyuan-dense` GGUFs whose `rope.freq_base` exceeds `1e6`
(Tencent's pipeline bakes the dynamic-NTK exponent into a static value),
the original HF config is restored: `rope_type="dynamic"`,
`rope_theta=10000`, `alpha=1000`. Otherwise long prompts diverge.
- New GGUF arch mapping `hunyuan-dense → hunyuan_v1_dense` with
`attn_q_norm`/`attn_k_norm` tensor name mappings.
- `_detect_quant_params` returns `block_size` explicitly so the
`QuantizationConfig` matches the actual per-type group size.

## Verification

End-to-end on `tencent/Hy-MT1.5-1.8B-2bit`:
- **Per-tensor dequantization** matches HF safetensors to bf16-rounding
precision (max abs `2.3e-5`).
- **Single-layer ORT matmul** vs HF dequantized matmul: max abs
`2.5e-3`, mean `2.8e-4`.
- **Short-prompt greedy generation** with the chat template produces
`Hello, world!<EOS>` token-for-token identical to HF.
- **Throughput** with default flag: 0.13 s prefill, 32.8 tok/s decode on
CPU EP.

Tests:
- Split into `TestTencentQ10DefaultInflated4Bit`,
`TestTencentQ10NativeBits2`, and `TestTencentQ10Shared` covering both
flag values + shape/error invariants.
- Full test suite (`tests/build_graph_test.py src/`): **2752 passed, 0
failures**.

## Out of scope / follow-ups

- 1.25-bit Sherry (`AngelSlim/Hy-MT1.5-1.8B-1.25bit-GGUF`, [llama.cpp PR
#22836](ggml-org/llama.cpp#22836) `STQ1_0`):
needs a ternary-with-structured-sparsity codebook that doesn't map
cleanly to `MatMulNBits`. Tracked in
[microsoft/onnxruntime#28549](microsoft/onnxruntime#28549).
- ORT MLAS fast path for `bits=2` + float zp:
[microsoft/onnxruntime#28552](microsoft/onnxruntime#28552).
When that lands, flip the flag default and remove the bits=4 inflation
path.
- ORT silent zero output for `block_size > 256`:
[microsoft/onnxruntime#28551](microsoft/onnxruntime#28551).

---------

Signed-off-by: Justin Chu <justinchu@microsoft.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <223894421+github-code-quality[bot]@users.noreply.github.com>
@wmm7777

wmm7777 commented Jul 6, 2026

Copy link
Copy Markdown

@sjl623 你好,实测 D8300s上测试,STQ1_0 prefill 速度(41 t/s)仅为 Q4_K_M(79 t/s)的一半, decode 速度(27 t/s)超过 Q4_K_M(22 t/s),感觉相较于Q4 decode 的优势不明显,prefill的差距很大,有优化计划吗

@Green-Sky

Copy link
Copy Markdown
Collaborator

This needs to be rebased and the ggml type id needs adjusting.

@sjl623

sjl623 commented Jul 20, 2026

Copy link
Copy Markdown
Author

@sjl623 你好,实测 D8300s上测试,STQ1_0 prefill 速度(41 t/s)仅为 Q4_K_M(79 t/s)的一半, decode 速度(27 t/s)超过 Q4_K_M(22 t/s),感觉相较于Q4 decode 的优势不明显,prefill的差距很大,有优化计划吗

STQ1_0 performs similarly to the existing ternary formats(TQ1_0, TQ2_0). Q4 benefits from mature tiled GEMM kernels, while ternary formats currently only have vec-dot kernels. Dedicated GEMM/GEMV optimizations may be explored in the future.

@sjl623

sjl623 commented Jul 20, 2026

Copy link
Copy Markdown
Author

This needs to be rebased and the ggml type id needs adjusting.

Done.

xuchen-intel added a commit to xuchen-intel/llama.cpp that referenced this pull request Jul 28, 2026
@892481092

Copy link
Copy Markdown

~/llama.cpp$ cmake --build build --config Release
[ 1%] Building C object ggml/src/CMakeFiles/ggml-base.dir/ggml.c.o
[ 1%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml.cpp.o
[ 1%] Building C object ggml/src/CMakeFiles/ggml-base.dir/ggml-alloc.c.o
[ 1%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-backend.cpp.o
[ 1%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-backend-meta.cpp.o
[ 2%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-opt.cpp.o
[ 2%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-threading.cpp.o
[ 2%] Building C object ggml/src/CMakeFiles/ggml-base.dir/ggml-quants.c.o
[ 2%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/gguf.cpp.o
[ 2%] Linking CXX shared library ../../bin/libggml-base.so
[ 2%] Built target ggml-base
[ 3%] Building C object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/ggml-cpu.c.o
[ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/ggml-cpu.cpp.o
[ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/repack.cpp.o
/home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp: In function ‘int repack_q4_0_to_q4_0_4_bl(ggml_tensor*, int, const void*, size_t)’:
/home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:2762:19: warning: writing 32 bytes into a region of size 0 [-Wstringop-overflow=]
2762 | memcpy(&out.qs[dst_offset], &elems, sizeof(uint64_t));
| ^
/home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:3222:20: note: at offset 72 into destination object ‘’ of size 72
3222 | *dst++ = make_block_q4_0x4(dst_tmp, interleave_block);
| ~~~~~~~^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
/home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:2762:19: warning: writing 32 bytes into a region of size 0 [-Wstringop-overflow=]
2762 | memcpy(&out.qs[dst_offset], &elems, sizeof(uint64_t));
| ^
/home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:3222:20: note: at offset 104 into destination object ‘’ of size 72
3222 | *dst++ = make_block_q4_0x4(dst_tmp, interleave_block);
| ~~~~~~~^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
[ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/hbm.cpp.o
[ 3%] Building C object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/quants.c.o
[ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/traits.cpp.o
[ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/amx/amx.cpp.o
[ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/amx/mmq.cpp.o
[ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/binary-ops.cpp.o
[ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/unary-ops.cpp.o
[ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/vec.cpp.o
[ 5%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/ops.cpp.o
[ 5%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/llamafile/sgemm.cpp.o
[ 5%] Building C object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/quants.c.o
[ 5%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/repack.cpp.o
[ 5%] Linking CXX shared library ../../bin/libggml-cpu.so
/usr/bin/ld: CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/quants.c.o: in function ggml_vec_dot_nvfp4_q8_0': quants.c:(.text+0xcb0): multiple definition of ggml_vec_dot_nvfp4_q8_0'; CMakeFiles/ggml-cpu.dir/ggml-cpu/quants.c.o:quants.c:(.text+0x1190): first defined here
collect2: error: ld returned 1 exit status
gmake[2]: *** [ggml/src/CMakeFiles/ggml-cpu.dir/build.make:327:bin/libggml-cpu.so.0.17.0] 错误 1
gmake[1]: *** [CMakeFiles/Makefile2:2409:ggml/src/CMakeFiles/ggml-cpu.dir/all] 错误 2
gmake: *** [Makefile:146:all] 错误 2

complie with an error

@yuziwe

yuziwe commented Aug 10, 2026

Copy link
Copy Markdown

~/llama.cpp$ cmake --build build --config Release [ 1%] Building C object ggml/src/CMakeFiles/ggml-base.dir/ggml.c.o [ 1%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml.cpp.o [ 1%] Building C object ggml/src/CMakeFiles/ggml-base.dir/ggml-alloc.c.o [ 1%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-backend.cpp.o [ 1%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-backend-meta.cpp.o [ 2%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-opt.cpp.o [ 2%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-threading.cpp.o [ 2%] Building C object ggml/src/CMakeFiles/ggml-base.dir/ggml-quants.c.o [ 2%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/gguf.cpp.o [ 2%] Linking CXX shared library ../../bin/libggml-base.so [ 2%] Built target ggml-base [ 3%] Building C object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/ggml-cpu.c.o [ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/ggml-cpu.cpp.o [ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/repack.cpp.o /home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp: In function ‘int repack_q4_0_to_q4_0_4_bl(ggml_tensor*, int, const void*, size_t)’: /home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:2762:19: warning: writing 32 bytes into a region of size 0 [-Wstringop-overflow=] 2762 | memcpy(&out.qs[dst_offset], &elems, sizeof(uint64_t)); | ^ /home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:3222:20: note: at offset 72 into destination object ‘’ of size 72 3222 | *dst++ = make_block_q4_0x4(dst_tmp, interleave_block); | ~~~~~~~^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ /home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:2762:19: warning: writing 32 bytes into a region of size 0 [-Wstringop-overflow=] 2762 | memcpy(&out.qs[dst_offset], &elems, sizeof(uint64_t)); | ^ /home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:3222:20: note: at offset 104 into destination object ‘’ of size 72 3222 | *dst++ = make_block_q4_0x4(dst_tmp, interleave_block); | ~~~~~~~^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ [ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/hbm.cpp.o [ 3%] Building C object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/quants.c.o [ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/traits.cpp.o [ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/amx/amx.cpp.o [ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/amx/mmq.cpp.o [ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/binary-ops.cpp.o [ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/unary-ops.cpp.o [ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/vec.cpp.o [ 5%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/ops.cpp.o [ 5%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/llamafile/sgemm.cpp.o [ 5%] Building C object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/quants.c.o [ 5%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/repack.cpp.o [ 5%] Linking CXX shared library ../../bin/libggml-cpu.so /usr/bin/ld: CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/quants.c.o: in function ggml_vec_dot_nvfp4_q8_0': quants.c:(.text+0xcb0): multiple definition of ggml_vec_dot_nvfp4_q8_0'; CMakeFiles/ggml-cpu.dir/ggml-cpu/quants.c.o:quants.c:(.text+0x1190): first defined here collect2: error: ld returned 1 exit status gmake[2]: *** [ggml/src/CMakeFiles/ggml-cpu.dir/build.make:327:bin/libggml-cpu.so.0.17.0] 错误 1 gmake[1]: *** [CMakeFiles/Makefile2:2409:ggml/src/CMakeFiles/ggml-cpu.dir/all] 错误 2 gmake: *** [Makefile:146:all] 错误 2

complie with an error

I also encountered this problem, solved by removing one marco define.

image

jinlongsong added 3 commits August 10, 2026 12:45
1.3125 bpw quantization. Each block of 256 elements stores 64 groups of 4
ternary lanes (-1/0/+1) with the constraint that every group has exactly
one zero and three non-zero lanes of identical magnitude. The ternary
pattern is encoded as a 4-bit codebook index plus a 1-bit global sign,
yielding 32 patterns over 4 lanes (5 bits / 4 lanes = 1.25 bpw payload),
plus a per-block fp16 scale (0.0625 bpw) for 1.3125 bpw total.

Components:
- block_stq_0 layout and codebook in ggml-common.h
- reference quantize/dequantize/validate in ggml-quants.{h,c}
- generic CPU vec_dot in ggml-cpu/quants.{h,c}
- ARM NEON vec_dot using vqtbl2q for codebook lookup, vdotq_s32 for
  accumulation, plus vld4q-based in-place repack of Q8_K activations
- enum slots: GGML_TYPE_STQ_0 = 42, LLAMA_FTYPE_MOSTLY_STQ_0 = 41,
  GGMLQuantizationType.STQ_0 = 42, LlamaFileType.MOSTLY_STQ_0 = 41
- llama-quantize CLI option "STQ_0"
@sjl623

sjl623 commented Aug 10, 2026

Copy link
Copy Markdown
Author

~/llama.cpp$ cmake --build build --config Release [ 1%] Building C object ggml/src/CMakeFiles/ggml-base.dir/ggml.c.o [ 1%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml.cpp.o [ 1%] Building C object ggml/src/CMakeFiles/ggml-base.dir/ggml-alloc.c.o [ 1%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-backend.cpp.o [ 1%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-backend-meta.cpp.o [ 2%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-opt.cpp.o [ 2%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/ggml-threading.cpp.o [ 2%] Building C object ggml/src/CMakeFiles/ggml-base.dir/ggml-quants.c.o [ 2%] Building CXX object ggml/src/CMakeFiles/ggml-base.dir/gguf.cpp.o [ 2%] Linking CXX shared library ../../bin/libggml-base.so [ 2%] Built target ggml-base [ 3%] Building C object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/ggml-cpu.c.o [ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/ggml-cpu.cpp.o [ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/repack.cpp.o /home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp: In function ‘int repack_q4_0_to_q4_0_4_bl(ggml_tensor*, int, const void*, size_t)’: /home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:2762:19: warning: writing 32 bytes into a region of size 0 [-Wstringop-overflow=] 2762 | memcpy(&out.qs[dst_offset], &elems, sizeof(uint64_t)); | ^ /home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:3222:20: note: at offset 72 into destination object ‘’ of size 72 3222 | *dst++ = make_block_q4_0x4(dst_tmp, interleave_block); | ~~~~~~~^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ /home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:2762:19: warning: writing 32 bytes into a region of size 0 [-Wstringop-overflow=] 2762 | memcpy(&out.qs[dst_offset], &elems, sizeof(uint64_t)); | ^ /home/liaoziyu/llama.cpp/ggml/src/ggml-cpu/repack.cpp:3222:20: note: at offset 104 into destination object ‘’ of size 72 3222 | *dst++ = make_block_q4_0x4(dst_tmp, interleave_block); | ~~~~~~~^~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ [ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/hbm.cpp.o [ 3%] Building C object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/quants.c.o [ 3%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/traits.cpp.o [ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/amx/amx.cpp.o [ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/amx/mmq.cpp.o [ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/binary-ops.cpp.o [ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/unary-ops.cpp.o [ 4%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/vec.cpp.o [ 5%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/ops.cpp.o [ 5%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/llamafile/sgemm.cpp.o [ 5%] Building C object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/quants.c.o [ 5%] Building CXX object ggml/src/CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/repack.cpp.o [ 5%] Linking CXX shared library ../../bin/libggml-cpu.so /usr/bin/ld: CMakeFiles/ggml-cpu.dir/ggml-cpu/arch/x86/quants.c.o: in function ggml_vec_dot_nvfp4_q8_0': quants.c:(.text+0xcb0): multiple definition of ggml_vec_dot_nvfp4_q8_0'; CMakeFiles/ggml-cpu.dir/ggml-cpu/quants.c.o:quants.c:(.text+0x1190): first defined here collect2: error: ld returned 1 exit status gmake[2]: *** [ggml/src/CMakeFiles/ggml-cpu.dir/build.make:327:bin/libggml-cpu.so.0.17.0] 错误 1 gmake[1]: *** [CMakeFiles/Makefile2:2409:ggml/src/CMakeFiles/ggml-cpu.dir/all] 错误 2 gmake: *** [Makefile:146:all] 错误 2
complie with an error

I also encountered this problem, solved by removing one marco define.

image

Thanks for feedback. The issues has fixed by rebase to latest master.

layfhaker added a commit to layfhaker/llama.cpp that referenced this pull request Aug 19, 2026
x86 port of the ARM NEON STQ1_0 kernel from Tencent Hunyuan PR
ggml-org#22836. AVX2 codebook lookup + maddubs, generic
fallback when AVX2 is unavailable.
@Green-Sky

Copy link
Copy Markdown
Collaborator

@BlackDawnNova

BlackDawnNova commented Aug 30, 2026 •

Copy link
Copy Markdown

Hi @eauchs — great PR. We tested STQ1_0 locally (Hunyuan-MT 1.8B and a custom
quant of Qwen3.8-Flash-Next) and have three findings, two of them with ready
patches:

(Also noticed @Green-Sky's pointer to the new Hy4-preview-GGUF — its STQ1_0
tensors will hit the same scalar-fallback wall on x86, so the kernel below is
becoming more relevant, not less.)

1. x86 AVX2 vec_dot kernel (the merge blocker) — patch ready

The PR currently only ships ARM NEON + scalar fallback, so x86 builds fall back
to the generic kernel (5.5 t/s vs 36.9 t/s for Q4_K_M on a Hunyuan-MT 1.8B test
model). We implemented the missing AVX2 kernel (sign-expanded codebook select via
pshufb+blendv, four 2-bit lane planes matched against contiguous stride-16 q8
slices, (q-1) correction from y block sums — mirroring the ARM version):

  • Result: Hunyuan-MT 1.8B STQ1_0 goes from 5.5 t/s (scalar) to 34.2 t/s
    (8 threads, 9800X3D), matching Q8_0-speed territory at 1.31 bpw.
  • Patch: [0001] below (98 lines), mirrors the ARM kernel structure 1:1.

2. Reference encoder is not SSE-optimal — patch ready

quantize_row_stq1_0_ref uses d = amax. With the zero assignment fixed, the
block error is sum(|x| - d)^2 over non-zero lanes, whose closed-form minimizer
is the mean |x| of those lanes, not amax. Grid-searching between mean and
amax per block:

  • Round-trip RMSE: 0.434 -> 0.261 (-40%)
  • Real impact: naive amax quantization of an untrained model produces unusable
    output (degenerate token loops, PPL ~1e28); with scale search the same model
    becomes coherent. Users WILL try this (we did), so we'd strongly recommend
    landing scale search with the reference encoder.
  • Patch: [0002] below (48 lines), format-compatible with existing kernels.

3. Pre-rename official GGUFs fail to load — enum mismatch (older files only)

The GGUFs published at AngelSlim/Hy-MT1.5-1.8B-1.25bit-GGUF (April/May) encode all
ternary weights with GGML_TYPE enum value 42, while this PR assigns
GGML_TYPE_STQ1_0 = 43 (the adjustment made on 2026-07-20 per @Green-Sky's
review). Anyone loading those older files with this PR's build gets a generic
"failed to load model" error. Verified by parsing the header (224 tensors at type
42); after rewriting the type fields to 43 the model loads and runs correctly on
this PR's kernel.

The new Hy4-preview-STQ1_0.gguf already uses 43 throughout, so this only affects
the pre-rename files — but since they are still up on HF and are the models most
people will try first (1.8B is a manageable download), a note in the PR
description or a loader-side shim would save every early adopter one very
confusing hour. (We hit it and confirmed via a header rewrite patch.)

Validation

  • llama-bench + llama-perplexity on Hunyuan-MT 1.8B (STQ1_0 vs Q8_0 baseline:
    PPL parity on the test set) and Qwen3.8-Flash-Next;
  • unit harness comparing AVX2 vs generic per-plane (4096 blocks, exact match);
  • kernel mirrors ARM semantics; verified ggml-cpu unit test included in the patch.

Both patches are against 1e411d8 on this PR's branch — happy to split/rebase
however you prefer.

Patch 0001 — ggml-cpu: AVX2 vec_dot kernel for STQ1_0 (98 lines)
From 994c9a086b58db739fbb7d2b24a5b776535fd498 Mon Sep 17 00:00:00 2001
From: ForiLusa <4963634@qq.com>
Date: Sun, 30 Aug 2026 02:03:13 +0800
Subject: [PATCH 1/2] ggml-cpu: AVX2 vec_dot kernel for STQ1_0

Native AVX2 implementation of ggml_vec_dot_stq1_0_q8_K for x86:
sign-expanded codebook select via pshufb+blendv, four 2-bit lane
planes matched against contiguous stride-16 q8 slices, (q-1)
correction from y block sums.

Fixes vs earlier WIP draft:
- half 1 groups (32..63) must OR sign bytes 4..7, not 0..3
- final reduction must cvtepi32_ps the int32 accumulator;
  castsi256_ps reinterprets bit patterns and turns negative
  partial sums into NaNs

Hunyuan-MT 1.8B STQ1_0: coherent output at ~33 t/s (16-thread
Zen4-ish desktop), PPL parity with q8_0 on the test set;
scalar fallback was 5.5 t/s.
---
 ggml/src/ggml-cpu/arch-fallback.h   |  1 -
 ggml/src/ggml-cpu/arch/x86/quants.c | 98 +++++++++++++++++++++++++++++
 2 files changed, 98 insertions(+), 1 deletion(-)

diff --git a/ggml/src/ggml-cpu/arch-fallback.h b/ggml/src/ggml-cpu/arch-fallback.h
index 259140f..eea39ec 100644
--- a/ggml/src/ggml-cpu/arch-fallback.h
+++ b/ggml/src/ggml-cpu/arch-fallback.h
@@ -85,7 +85,6 @@
 #elif defined(__x86_64__) || defined(__i386__) || defined(_M_IX86) || defined(_M_X64)
 // quants.c
 #define ggml_vec_dot_q2_0_q8_0_generic ggml_vec_dot_q2_0_q8_0
-#define ggml_vec_dot_stq1_0_q8_K_generic ggml_vec_dot_stq1_0_q8_K
 // repack.cpp
 #define ggml_quantize_mat_q8_0_4x4_generic ggml_quantize_mat_q8_0_4x4
 #define ggml_quantize_mat_q8_K_4x4_generic ggml_quantize_mat_q8_K_4x4
diff --git a/ggml/src/ggml-cpu/arch/x86/quants.c b/ggml/src/ggml-cpu/arch/x86/quants.c
index ea54cfe..5bc85b7 100644
--- a/ggml/src/ggml-cpu/arch/x86/quants.c
+++ b/ggml/src/ggml-cpu/arch/x86/quants.c
@@ -1763,6 +1763,104 @@ void ggml_vec_dot_q2_K_q8_K(int n, float * GGML_RESTRICT s, size_t bs, const voi
 #endif
 }
 
+void ggml_vec_dot_stq1_0_q8_K(int n, float * GGML_RESTRICT s, size_t bs, const void * GGML_RESTRICT vx, size_t bx, const void * GGML_RESTRICT vy, size_t by, int nrc) {
+    assert(nrc == 1);
+    UNUSED(nrc);
+    UNUSED(bx);
+    UNUSED(by);
+    UNUSED(bs);
+
+    const block_stq1_0 * GGML_RESTRICT x = vx;
+    const block_q8_K  * GGML_RESTRICT y = vy;
+
+    const int nb = n / QK_K;
+
+#if defined(__AVX2__)
+    const __m128i m0f = _mm_set1_epi8(0x0F);
+    const __m128i m03 = _mm_set1_epi8(0x03);
+    const __m128i v16 = _mm_set1_epi8(16);
+    const __m128i cb_lo = _mm_loadu_si128((const __m128i *) stq1_0_codebook);
+    const __m128i cb_hi = _mm_loadu_si128((const __m128i *) (stq1_0_codebook + 16));
+    const __m256i ones16 = _mm256_set1_epi16(1);
+
+    float sumf = 0.0f;
+    for (int i = 0; i < nb; ++i) {
+        __m256i acc = _mm256_setzero_si256();
+
+        const uint8_t * sg = x[i].sign;
+        // expand the 8 sign bytes to 64 group sign bytes (0x10 / 0x00)
+        uint8_t sgn[64];
+        for (int j = 0; j < 64; ++j) {
+            sgn[j] = ((sg[j >> 3] >> (j & 7)) & 1) ? 0x10 : 0x00;
+        }
+
+        const __m128i qs0 = _mm_loadu_si128((const __m128i *) (x[i].qs));
+        const __m128i qs1 = _mm_loadu_si128((const __m128i *) (x[i].qs + 16));
+
+        // two 16-group chunks: chunk h covers groups i*32+16h..+15 and y bytes [i*128+64h .. +63]
+        for (int h = 0; h < 2; ++h) {
+            // sign select bytes for this half: groups 32h..32h+31 = sign bytes 4h..4h+3
+            const __m128i s0 = _mm_loadu_si128((const __m128i *) (sgn + 32*h));
+            const __m128i s1 = _mm_loadu_si128((const __m128i *) (sgn + 32*h + 16));
+            const __m128i packed = (h == 0) ? qs0 : qs1;
+            const __m128i lo = _mm_and_si128(packed, m0f);
+            const __m128i hi = _mm_and_si128(_mm_srli_epi16(packed, 4), m0f);
+            const __m128i ev = _mm_unpacklo_epi8(lo, hi);   // groups h*32+0..7 (code+sign OR'd)
+            const __m128i od = _mm_unpackhi_epi8(lo, hi);   // groups h*32+8..15
+
+            const __m128i evs = _mm_or_si128(ev, s0);
+            const __m128i ods = _mm_or_si128(od, s1);
+            // select the codebook half by the sign bit: 0x80 bytes take cb_hi, others cb_lo
+            const __m128i msk_e = _mm_slli_epi16(_mm_and_si128(evs, _mm_set1_epi8(0x10)), 3);
+            const __m128i msk_o = _mm_slli_epi16(_mm_and_si128(ods, _mm_set1_epi8(0x10)), 3);
+            const __m128i sel_e = _mm_blendv_epi8(_mm_shuffle_epi8(cb_lo, evs),
+                                                  _mm_shuffle_epi8(cb_hi, _mm_sub_epi8(evs, v16)), msk_e);
+            const __m128i sel_o = _mm_blendv_epi8(_mm_shuffle_epi8(cb_lo, ods),
+                                                  _mm_shuffle_epi8(cb_hi, _mm_sub_epi8(ods, v16)), msk_o);
+            const int8_t * q8 = y[i].qs + 128*h;
+            acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+                    _mm256_cvtepi8_epi16(_mm_and_si128(sel_e, m03)),
+                    _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 +  0)))));
+            acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+                    _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_e, 2), m03)),
+                    _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 16)))));
+            acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+                    _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_e, 4), m03)),
+                    _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 32)))));
+            acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+                    _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_e, 6), m03)),
+                    _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 48)))));
+            acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+                    _mm256_cvtepi8_epi16(_mm_and_si128(sel_o, m03)),
+                    _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 64)))));
+            acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+                    _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_o, 2), m03)),
+                    _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 80)))));
+            acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+                    _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_o, 4), m03)),
+                    _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 96)))));
+            acc = _mm256_add_epi32(acc, _mm256_madd_epi16(
+                    _mm256_cvtepi8_epi16(_mm_and_si128(_mm_srli_epi16(sel_o, 6), m03)),
+                    _mm256_cvtepi8_epi16(_mm_loadu_si128((const __m128i *) (q8 + 112)))));
+        }
+
+        // (q-1) correction: subtract the y sum over the block's 256 positions
+        const __m256i ys = _mm256_loadu_si256((const __m256i *) y[i].bsums);
+        acc = _mm256_sub_epi32(acc, _mm256_madd_epi16(ys, _mm256_set1_epi16(1)));
+        // cvtepi32_ps converts the int32 lanes; castsi256_ps would reinterpret the
+        // bit patterns and turn negative partial sums into NaNs
+        sumf += (GGML_CPU_FP16_TO_FP32(x[i].d) * y[i].d) * hsum_float_8(_mm256_cvtepi32_ps(acc));
+    }
+
+    *s = sumf;
+#else
+    UNUSED(x);
+    UNUSED(y);
+    UNUSED(nb);
+    ggml_vec_dot_stq1_0_q8_K_generic(n, s, bs, vx, bx, vy, by, nrc);
+#endif
+}
+
 void ggml_vec_dot_q3_K_q8_K(int n, float * GGML_RESTRICT s, size_t bs, const void * GGML_RESTRICT vx, size_t bx, const void * GGML_RESTRICT vy, size_t by, int nrc) {
     assert(n % QK_K == 0);
     assert(nrc == 1);
-- 
2.54.0.windows.1

Patch 0002 — quantize: SSE-optimal scale search for STQ1_0 (48 lines)
From 01bad5d3247d3b04b6c467e340f95c707f6fccae Mon Sep 17 00:00:00 2001
From: ForiLusa <4963634@qq.com>
Date: Sun, 30 Aug 2026 11:55:29 +0800
Subject: [PATCH 2/2] quantize: SSE-optimal scale search for STQ1_0

With the per-group zero assignment fixed, the block reconstruction
error is sum(|x| - d)^2 over non-zero lanes, whose closed-form
minimizer is the mean |x| of those lanes - not amax. Grid-search
between the mean and amax per block and keep the lowest-SSE scale.

Format-compatible with existing kernels. Nemotron-H-8B f16->STQ1_0
round-trip RMSE 0.434 -> 0.261; wiki PPL (ffn-only ternary mix,
rest q8_0) 45307 -> 83.
---
 ggml/src/ggml-quants.c | 50 ++++++++++++++++++++++++++++++++++++++++--
 1 file changed, 48 insertions(+), 2 deletions(-)

diff --git a/ggml/src/ggml-quants.c b/ggml/src/ggml-quants.c
index 00ff915..d0c2ba7 100644
--- a/ggml/src/ggml-quants.c
+++ b/ggml/src/ggml-quants.c
@@ -2424,12 +2424,58 @@ void quantize_row_stq1_0_ref(const float * GGML_RESTRICT x, block_stq1_0 * GGML_
             const float a = fabsf(x[j]);
             if (a > amax) amax = a;
         }
-        y[i].d = GGML_FP32_TO_FP16(amax);
 
         // STQ1_0 forces exactly one zero per group of 4. Groups are stride-16
         // within each 64-weight chunk: group g (chunk-local in 0..15) covers
         // {x[c*64 + g + p*16] : p in 0..3}. Pick the smallest-|x| lane as zero;
-        // project the other 3 onto {-d, +d} via sign.
+        // project the other 3 onto {-d, +d}.
+        //
+        // d = amax is not SSE-optimal: with the zero assignment fixed, the
+        // block error is sum(|x| - d)^2 over non-zero lanes, whose closed-form
+        // minimizer is the mean |x| of those lanes. Grid-search between the
+        // mean and amax and keep the candidate with the lowest block SSE.
+        float mean_abs  = 0.0f;
+        int   n_nonzero = 0;
+        for (int g = 0; g < QK_K/4; ++g) {
+            const int chunk = g / 16;
+            const int gloc  = g % 16;
+            const float * base = x + chunk*64 + gloc;
+            for (int p = 0; p < 4; ++p) {
+                mean_abs += fabsf(base[p*16]);
+            }
+            n_nonzero += 3;
+        }
+        mean_abs = (n_nonzero > 0) ? mean_abs / n_nonzero : amax;
+
+        float best_d   = amax;
+        float best_sse = 1e30f;
+        for (int t = 0; t <= 7; ++t) {
+            const float d = mean_abs + (amax - mean_abs) * (t / 7.0f);
+            float sse = 0.0f;
+            for (int g = 0; g < QK_K/4; ++g) {
+                const int chunk = g / 16;
+                const int gloc  = g % 16;
+                const float * base = x + chunk*64 + gloc;
+                int zero_pos = 0;
+                float min_abs = fabsf(base[0]);
+                for (int p = 1; p < 4; ++p) {
+                    const float a = fabsf(base[p*16]);
+                    if (a < min_abs) { min_abs = a; zero_pos = p; }
+                }
+                for (int p = 0; p < 4; ++p) {
+                    if (p == zero_pos) {
+                        sse += base[p*16] * base[p*16];
+                    } else {
+                        const float e = fabsf(base[p*16]) - d;
+                        sse += e * e;
+                    }
+                }
+            }
+            if (sse < best_sse) { best_sse = sse; best_d = d; }
+        }
+
+        y[i].d = GGML_FP32_TO_FP16(best_d);
+
         for (int g = 0; g < QK_K/4; ++g) {
             const int chunk = g / 16;
             const int gloc  = g % 16;
-- 
2.54.0.windows.1

Sube-py pushed a commit to Sube-py/llama.cpp that referenced this pull request Aug 31, 2026
STQ1_0 (coming via ggml-org#22836) has no Metal kernel. Without this guard,
ggml_metal_supports_mul_mat_op returns true for it, the scheduler
places STQ1_0 weights in Metal buffers when -ngl > 0, and compute
aborts with 'Asserting on type 43' / EXC_BAD_ACCESS in the
pipeline lookup.

Reserves the GGML_TYPE_STQ1_0 enum id (43) so it matches the
value the STQ PR uses. Once ggml-org#22836 merges, both branches agree.
@BlackDawnNova

Copy link
Copy Markdown

The Hy4 merge made this more urgent: mainline still has no STQ1_0, so the AngelSlim Hy4 STQ1_0 GGUFs fail on stock builds ("invalid ggml type 43" — saw a user report exactly this a few days ago).

The two patches I posted above (AVX2 vec_dot kernel + SSE scale search) still apply to the current head, re-checked today. Numbers from when I wrote them, Hunyuan-MT 1.8B, 8 threads on a 9800X3D: 5.5 t/s with the scalar fallback, 34.2 t/s with the kernel. That's Q8_0 territory at 1.31 bpw.

I can keep the patches rebased on this branch and handle the x86 side in review, so missing x86 support isn't a reason to hold the merge. For anyone blocked on type 43 right now: it's only the type-field enum in the GGUFs (42 vs 43), rewriting those fields makes the model load fine. Small python script does it, ask if you need it.

@devYRPauli

Copy link
Copy Markdown
Contributor

I re-measured on master at 41abbfd59, today.

git grep STQ1_0 on master returns nothing. The type is still not in mainline.

The load failure has an exact cause. On master GGML_TYPE_COUNT is 43. ggml/src/gguf.cpp:715 rejects any tensor type outside [0, GGML_TYPE_COUNT). This PR assigns GGML_TYPE_STQ1_0 = 43. So the released GGUFs carry the id this PR adds, and a stock build rejects it by one slot.

On the rebase question: the head 1e411d8f5 is 633 commits behind master. git merge-tree upstream/master pr22836 reports 0 conflicts.

I merged it in a worktree, built on an M1 Pro with Metal off, and ran test-quantize-fns:

check result
absolute quantization error ok, 0.013784
reference implementation error ok, 0.000000
dot product error ok, 0.144828
total 0 tests failed

test-quantize-fns is the test that covers the CPU vec_dot, and this PR already registers STQ1_0 in both of its error bounds. It runs on the x86 runners. So the AVX2 kernel in @BlackDawnNova's patch gets checked there once this lands. I cannot run AVX2 here.

One thing that does not help. Master now enables GGML_TYPE_TQ1_0 and GGML_TYPE_TQ2_0 in all_types and other_types in tests/test-backend-ops.cpp. Adding STQ1_0 to those two lists does not test the CPU kernels. I added it, built, and ran test-backend-ops test -o MUL_MAT. The run prints Skipping CPU backend, because the CPU backend is the reference that the other backends are compared against.

@BlackDawnNova

Copy link
Copy Markdown

@devYRPauli thanks for doing this legwork, it's more than I had. Good to know test-quantize-fns runs on the x86 runners with STQ1_0 bounds already registered, so the AVX2 kernel from my patch gets CI coverage as soon as it lands.

Agreed on the backend-ops part. Adding STQ1_0 to those lists would just print Skipping CPU backend, since the CPU is the reference the others get compared against. Honest coverage for the CPU kernel today is the error-bounds test plus real-weights runs (that's where my 34 t/s number came from). If a proper cross-backend test shows up that treats the CPU as an implementation instead of the reference, I'll wire STQ1_0 into it.

And if anyone hits the type 43 wall on the old AngelSlim files, ping me here, the fix is a small field-rewrite script, I'll paste it.

Sube-py pushed a commit to Sube-py/llama.cpp that referenced this pull request Sep 17, 2026
STQ1_0 (coming via ggml-org#22836) has no Metal kernel. Without this guard,
ggml_metal_supports_mul_mat_op returns true for it, the scheduler
places STQ1_0 weights in Metal buffers when -ngl > 0, and compute
aborts with 'Asserting on type 43' / EXC_BAD_ACCESS in the
pipeline lookup.

Reserves the GGML_TYPE_STQ1_0 enum id (43) so it matches the
value the STQ PR uses. Once ggml-org#22836 merges, both branches agree.

Assisted-by: OpenAI

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion examples ggml changes relating to the ggml tensor library for machine learning python python script changes testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.