Repository navigation
Add Q1_0 GGUF support (mainline + Tencent SEQ 2-bit variant) - #302
Conversation
This adds two paths through the GGUF loader for GGML type 41:
1. **Mainline Q1_0** (1-bit binary, 18 bytes / 128 elements): inflated
to MatMulNBits bits=2 with zero_point=1 so the {-d, +d} codebook
becomes scale·(B - 1) for B ∈ {0, 2}. Verified against
prism-ml/Bonsai-1.7B-Q1_0 reference checkpoints.
2. **Tencent custom Q1_0** (2-bit SEQ, 130 bytes / 512 elements):
used by AngelSlim/Hy-MT1.5-1.8B-2bit-GGUF. The on-disk layout is
per-row `[fp16 scale][2-bit codes packed LSB-first]` with the
SEQ codebook {-1.5, -0.5, 0.5, 1.5} (stored scale = effective/2,
equivalently codebook {-3, -1, +1, +3}·scale). Mainline llama.cpp
refuses to load these files because every tensor after the first
Q1_0 entry lands at the wrong offset — we bypass gguf-py's size
calculation and read each tensor from its explicit per-tensor
offset.
The Tencent layout is inflated to MatMulNBits bits=4 with
zero_point=3 (codes ∈ {0,2,4,6}) so dequant gives
scale·(B - 3) ∈ {-3, -1, +1, +3}·scale exactly. We use bits=4
instead of native bits=2 because the ORT CPU kernel currently
only implements float zero_points for bits=4, and the SEQ offset
1.5 cannot be expressed with an integer zp. The native 512-element
block is replicated across four 128-element MatMulNBits sub-blocks
(block_size=512 silently returns zeros on CPU EP).
Other infrastructure changes:
- QuantizedLinear now accepts bits ∈ {2, 4, 8}; zero_point packing
generalized to `ceil(n_blocks * bits / 8)` bytes.
- GGUF→ArchitectureConfig conversion now derives `rope_type` from
`<arch>.rope.scaling.type` (defaulting to "default" when
`rope.freq_base` is present), fixing models whose graphs were
previously built without RoPE.
- Added `hunyuan-dense` → `hunyuan_v1_dense` GGUF arch mapping and
the corresponding attn_q_norm/attn_k_norm tensor name mappings.
- `_detect_quant_params` now returns block_size explicitly so the
QuantizationConfig matches the actual per-type group size.
End-to-end verified on tencent/Hy-MT1.5-1.8B-2bit: greedy generation
produces `Hello, world!<EOS>` token-for-token identical to HF, and
per-tensor dequantization matches HF safetensors to bf16-rounding
precision (max abs 2.3e-5).
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
Performance Comparison
|
There was a problem hiding this comment.
Pull request overview
Adds GGUF loader support for GGML_TYPE_Q1_0 (id 41) in two flavors: mainline llama.cpp 1-bit binary (used by prism-ml/Bonsai) and Tencent's custom 2-bit SEQ variant used by AngelSlim/Hy-MT1.5-1.8B-2bit-GGUF. Mainline Q1_0 is repacked into ORT MatMulNBits with bits=2, zp=1; Tencent Q1_0 is read via an offset-based parser and inflated into bits=4, zp=3 to work around ORT CPU kernel limitations on 2-bit float zero-points and 512-element block sizes. Several supporting infrastructure changes (generalized QuantizedLinear for bits=2, rope_type derivation from GGUF rope.scaling.type, hunyuan-dense arch mapping, and block_size propagation from _detect_quant_params) round out the change.
Changes:
- Add Q1_0 (mainline 1-bit) repack path in
_repacker.pyplus extensive tests. - Add custom Tencent Q1_0 parser (
_tencent_q1_0.py) that reads tensors from explicit GGUF offsets and produces 4-bitMatMulNBitsarrays, integrated into_builder.py. - Generalize
QuantizedLinearto supportbits ∈ {2,4,8}, deriverope_typefor GGUF, registerhunyuan-dense, and surfaceblock_sizefrom quant detection.
Reviewed changes
Copilot reviewed 10 out of 10 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| src/mobius/integrations/gguf/_tensor_mapping.py | Register hunyuan-dense/hunyuan_v1_dense mapping and add attn_q_norm/attn_k_norm Hunyuan extras. |
| src/mobius/integrations/gguf/_tencent_q1_0.py | New Tencent Q1_0 layout detector and per-tensor parser producing MatMulNBits bits=4, zp=3 arrays. |
| src/mobius/integrations/gguf/_tencent_q1_0_test.py | Unit tests for the Tencent parser: shape, scale replication, bit packing, and dequant round-trip. |
| src/mobius/integrations/gguf/_repacker.py | Add mainline Q1_0 (1-bit) repack path → MatMulNBits 2-bit with shared zp=1. |
| src/mobius/integrations/gguf/_repacker_test.py | New Q1_0 repack tests covering shape, packing, bit order, and round-trip dequant. |
| src/mobius/integrations/gguf/_config_mapping.py | Derive rope_type from <arch>.rope.scaling.type and register Hunyuan model_type. |
| src/mobius/integrations/gguf/_builder.py | Return block_size from quant detection, plumb Tencent Q1_0 path through quantized state-dict loading. |
| src/mobius/integrations/gguf/_builder_test.py | Update _detect_quant_params signature usage and check block_size. |
| src/mobius/components/_quantized_linear.py | Allow bits=2; generalize zero-point packing to ceil(n_blocks * bits / 8). |
| src/mobius/components/_quantized_linear_test.py | Add tests for bits=2 shapes, zero-point packing, and forward attributes. |
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
🏗️ Architecture Diff
No architecture changes detected. ✅ Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed) |
Co-authored-by: Copilot Autofix powered by AI <223894421+github-code-quality[bot]@users.noreply.github.com> Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Earlier wording implied MatMulNBits doesn't support bits=2 at all.
Spec actually allows bits in {2..8}, and bits=2 with integer
zero_points works on the CPU EP. The two issues that actually force
the bits=4 / block_size=128 workarounds for SEQ are:
1. The unpacked (float) zero_point CPU path is only implemented
for bits=4 — needed because SEQ's offset 1.5 isn't an integer.
2. bits=4, block_size=512 silently returns zeros on CPU EP.
Both have follow-up ORT issues. Updated module + builder docstrings
to reflect this precisely.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
Replace the bits=4 inflation workaround with the native 2-bit
representation. The on-disk weight bytes go through unchanged
(LSB-first 4-codes-per-byte matches MatMulNBits), and the SEQ
codebook { -1.5, -0.5, +0.5, +1.5 } * stored_scale is expressed as
effective_scale = 2 * stored_scale
zero_point = 1.5
dequant = effective_scale * (code - 1.5)
= stored_scale * {-3, -1, +1, +3}[code]
This relies on the bits=2 float zero-point CPU fallback added in
microsoft/onnxruntime#28354 (merged 2026-05-13, shipping in ORT
1.27+). Single-layer parity verified: dequantized matmul against
the HF safetensors weight matches to max_abs=2.5e-3 (bf16 rounding).
Short-prompt greedy generation reproduces HF's translation
token-for-token through EOS.
Infrastructure changes:
- QuantizedLinear gains a zero_point_dtype argument; UINT8 keeps
the bit-packed integer form, float dtypes give one un-packed
value per block.
- QuantizationConfig gains float_zero_point; TextModel passes
config.dtype as zero_point_dtype when it is set.
- GGUF builder sets float_zero_point=True for Tencent Q1_0
layouts (bits=2, block_size=128 because ORT silently zeros at
block_size > 256 — see microsoft/onnxruntime#28551).
- GGUF -> ArchitectureConfig restores HF dynamic-NTK RoPE
(rope_theta=10000, alpha=1000) for hunyuan-dense files where
Tencent baked a large freq_base (>1e6) into the metadata.
Without this, long prompts diverge because the dynamic exponent
isn't reapplied per position.
Test coverage updated to reflect the new tensor shapes:
bits=2, block_size=128, float32 zero_points = 1.5, scales =
2 * stored_scale.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
ORT 1.27's bits=2 + float-zp CPU kernel is a naive scalar fallback
("naive implementation, need to be optimized" per matmul_nbits.cc),
giving ~0.24 tok/s decode on this model vs ~32 tok/s for the bits=4
packed path -- ~130x slower on the same weights. Until an MLAS fused
kernel lands (tracked in microsoft/onnxruntime#28552), default to the
fast inflated representation and put the smaller native form behind
the MOBIUS_TENCENT_Q1_0_USE_NATIVE_2BIT flag.
Changes:
- New flag tencent_q1_0_use_native_2bit (default False).
- parse_tencent_q1_0_tensor splits into _read_tencent_blocks +
_pack_inflated_4bit / _pack_native_2bit; dispatches on flag.
- GGUF builder feeds float_zero_point=True only when the flag is on.
- _detect_quant_params returns (bits=4, block_size=128) by default
and (bits=2, block_size=128) when the flag is on.
- Test classes split: TestTencentQ10DefaultInflated4Bit for the
fast default, TestTencentQ10NativeBits2 for the flag-on path,
TestTencentQ10Shared for shape/error tests.
End-to-end on tencent/Hy-MT1.5-1.8B-2bit with default flag:
prefill 0.13s, decode 32.8 tok/s on CPU EP.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
- Replace ambiguous unicode (-, x, /, etc.) with ASCII equivalents in changed-file docstrings and comments (RUFF/RUF002, RUF003). - Rename uppercase N -> n in test locals (RUFF/N806). - Reflow native-path test docstring summary (RUFF/D205). - Apply ruff format auto-fixes across the PR diff. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
c7f9343 to
bce4293
Compare
ORT GenAI's LLM model_type registry (onnxruntime-genai src/models/model_type.h) does not include hunyuan_v1_dense, so the genai_config.json mobius wrote for the Hy-MT1.5-1.8B-2bit export failed to load. The registry does include 'decoder' as a generic catch-all for any decoder-only causal LM that follows the standard input/output shapes, which is exactly what hunyuan_v1_dense is. This patch adds the mapping and bundles a runnable example (examples/hy_mt1_5.py) that exercises the full pipeline: build (bf16 or Q1_0) -> ModelPackage.save -> write_ort_genai_config -> og.Model / og.Tokenizer / og.Generator End-to-end verified on /tmp/hymt-q1-genai (Q1_0 quantized export, CPU EP, ORT GenAI 0.13.2): session loads, generator runs, generate_next_token() returns logical tokens. The example also patches past_present_share_buffer to False post-hoc because mobius currently emits share_buffer=True for the CPU EP even though our exports use a dynamic-shape KV cache; that mismatch is documented inline as a separate follow-up. A known downstream caveat is documented in the example header: ort-extensions' BPE tokenizer does not round-trip the upstream Hy-MT1.5 vocab for CJK content (e.g. '你好' tokenizes to an empty sequence), so the model is fed an English-only prompt and replies 'Please provide the text you want to be translated.' The decoder loads and generates correctly; the tokenizer parity is independent of this PR. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Signed-off-by: Justin Chu <justinchu@microsoft.com>
bce4293 to
cb26b68
Compare
The example originally used og.Tokenizer, but the Hy-MT1.5 vocab's
custom regex pre-tokenizer is not currently round-tripped by
ort-extensions (e.g. 'Hello, world!' tokenizes to a single space
token, '你好' tokenizes to []), so the model never saw the actual
prompt content. By default we now tokenize with HF and feed raw
token IDs to og.Generator -- still exercises the full ORT GenAI
inference path. Pass --use-ort-tokenizer to reproduce the broken
path for debugging.
Verified working end-to-end on the Q1_0 quantized export, CPU EP:
'Translate to Spanish: The cat is sleeping on the chair.'
-> 'El gato está durmiendo en la silla.'
'Translate to French: Knowledge is power.'
-> 'La connaissance est puissance.'
'Translate the following Chinese text to English: 你好,世界!'
-> 'Hello, world!'
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
16162c5 to
fb59741
Compare
…-gguf
ORT GenAI's past_present_share_buffer mode requires the decoder graph
to update the KV cache in place (com.microsoft.GroupQueryAttention).
The standard ONNX opset 23 Attention op concatenates past_key + new_K
and returns a dynamic-shape present_key, which is incompatible with
the pre-allocated shared buffer ORT GenAI passes in. The previous
config wrote past_present_share_buffer=True purely from the target EP
flag, which meant builds that didn't apply the GQA rewrite (anything
other than --ep cpu / cuda / dml / etc.) silently shipped a broken
config and crashed at first inference with 'inconsistent
total_sequence_length'.
This change:
* Adds supports_in_place_kv_cache to GenaiConfigGenerator and
_default_search_params; when set it overrides the EP flag.
* In auto_export._write_genai_config, introspects the decoder graph
for GroupQueryAttention nodes and passes the result through.
Standard-Attention graphs now correctly get share_buffer=False
regardless of the target EP.
* Adds --ep to the build-gguf CLI and threads execution_provider
through build_from_gguf so users can opt into the GQA rewrite for
GGUF imports the same way they can for HF imports.
* Updates examples/hy_mt1_5.py to build the Q1_0 variant with
--ep cpu so the example exercises the full share-buffer path
end-to-end and removes the post-hoc _patch_genai_config workaround.
Verified end-to-end on tencent/Hy-MT1.5-1.8B-2bit:
- 32 GroupQueryAttention nodes after build
- genai_config.search.past_present_share_buffer = true
- 'Translate to Spanish: The cat is sleeping on the chair.'
-> 'El gato está durmiendo en la silla.'
Full suite: 2753 passed, 0 failures.
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
Summary
Adds two paths through the GGUF loader for GGML type 41 so mobius can ingest both standard and Tencent's custom variants of Q1_0.
1. Mainline Q1_0 — 1-bit binary
Layout (18 bytes / 128 elements):
[fp16 d][16B packed 1-bit signs]. Dequant:bit ? +d : -d. Used by prism-ml/Bonsai checkpoints.Mapped to
MatMulNBitsbits=2,zero_point=1, codes ∈ {0, 2}, scale=d.2. Tencent custom Q1_0 — 2-bit SEQ
Used by
AngelSlim/Hy-MT1.5-1.8B-2bit-GGUF. Per-row layout:[fp16 stored_scale][2-bit codes packed LSB-first], native block size 512, codebook{−3, −1, +1, +3} · stored_scale.Mainline llama.cpp refuses to load these files because every tensor after the first Q1_0 entry lands at the wrong offset. We bypass the size calculation by reading each tensor from its explicit GGUF offset.
Two
MatMulNBitsrepresentations, selected by a flag:MOBIUS_TENCENT_Q1_0_USE_NATIVE_2BIT=12c ∈ {0,2,4,6}, integer zp=3)Both representations produce bit-identical dequantized weights matching HF safetensors to bf16 rounding (
max_abs ≈ 2.5e-3vs HF reference matmul). The 130× decode speed difference is purely the ORT CPU kernel: the bits=4 path is the mature MLAS fused path, the bits=2 + float-zp path is the kernel's own self-described "naive implementation, need to be optimized" scalar fallback. Tracked in microsoft/onnxruntime#28552; once that lands the default can flip.Block-size workaround: ORT silently returns zeros for
block_size > 256(#28551). Both representations therefore exposeblock_size = 128by replicating each native 512-element scale across 4 sub-blocks.Other infrastructure changes
tencent_q1_0_use_native_2bit(defaultFalse).QuantizedLinearacceptsbits ∈ {2, 4, 8}and a newzero_point_dtypeargument.UINT8(default) keeps the bit-packed integer form; float dtypes produce one un-packed value per block.QuantizationConfiggainsfloat_zero_point;TextModelpassesconfig.dtypeas the zp dtype when it is set.ArchitectureConfigderivation now setsrope_typefrom<arch>.rope.scaling.type(defaulting to"default"whenrope.freq_baseis present). Without this fix, models withrope.scaling.type = "none"were built with no RoPE at all.hunyuan-denseGGUFs whoserope.freq_baseexceeds1e6(Tencent's pipeline bakes the dynamic-NTK exponent into a static value), the original HF config is restored:rope_type="dynamic",rope_theta=10000,alpha=1000. Otherwise long prompts diverge.hunyuan-dense → hunyuan_v1_densewithattn_q_norm/attn_k_normtensor name mappings._detect_quant_paramsreturnsblock_sizeexplicitly so theQuantizationConfigmatches the actual per-type group size.Verification
End-to-end on
tencent/Hy-MT1.5-1.8B-2bit:2.3e-5).2.5e-3, mean2.8e-4.Hello, world!<EOS>token-for-token identical to HF.Tests:
TestTencentQ10DefaultInflated4Bit,TestTencentQ10NativeBits2, andTestTencentQ10Sharedcovering both flag values + shape/error invariants.tests/build_graph_test.py src/): 2752 passed, 0 failures.Out of scope / follow-ups
AngelSlim/Hy-MT1.5-1.8B-1.25bit-GGUF, llama.cpp PR #22836STQ1_0): needs a ternary-with-structured-sparsity codebook that doesn't map cleanly toMatMulNBits. Tracked in microsoft/onnxruntime#28549.bits=2+ float zp: microsoft/onnxruntime#28552. When that lands, flip the flag default and remove the bits=4 inflation path.block_size > 256: microsoft/onnxruntime#28551.