Skip to content

Add Q1_0 GGUF support (mainline + Tencent SEQ 2-bit variant) - #302

Merged
justinchuby merged 10 commits into
mainfrom
add-q1-0-gguf-support
May 20, 2026
Merged

justinchuby merged 10 commits into
mainfrom
add-q1-0-gguf-support

Conversation

@justinchuby

@justinchuby justinchuby commented May 18, 2026 •

Copy link
Copy Markdown
Member

Summary

Adds two paths through the GGUF loader for GGML type 41 so mobius can ingest both standard and Tencent's custom variants of Q1_0.

1. Mainline Q1_0 — 1-bit binary

Layout (18 bytes / 128 elements): [fp16 d][16B packed 1-bit signs]. Dequant: bit ? +d : -d. Used by prism-ml/Bonsai checkpoints.

Mapped to MatMulNBits bits=2, zero_point=1, codes ∈ {0, 2}, scale=d.

2. Tencent custom Q1_0 — 2-bit SEQ

Used by AngelSlim/Hy-MT1.5-1.8B-2bit-GGUF. Per-row layout: [fp16 stored_scale][2-bit codes packed LSB-first], native block size 512, codebook {−3, −1, +1, +3} · stored_scale.

Mainline llama.cpp refuses to load these files because every tensor after the first Q1_0 entry lands at the wrong offset. We bypass the size calculation by reading each tensor from its explicit GGUF offset.

Two MatMulNBits representations, selected by a flag:

default MOBIUS_TENCENT_Q1_0_USE_NATIVE_2BIT=1
ORT bits 4 (inflated 2c ∈ {0,2,4,6}, integer zp=3) 2 (codes pass-through, float zp=1.5)
on-disk weight bpw 4 (2× the source) 2 (matches source)
CPU decode throughput ~32 tok/s ~0.24 tok/s
ORT version any ≥ 1.27 (#28354)

Both representations produce bit-identical dequantized weights matching HF safetensors to bf16 rounding (max_abs ≈ 2.5e-3 vs HF reference matmul). The 130× decode speed difference is purely the ORT CPU kernel: the bits=4 path is the mature MLAS fused path, the bits=2 + float-zp path is the kernel's own self-described "naive implementation, need to be optimized" scalar fallback. Tracked in microsoft/onnxruntime#28552; once that lands the default can flip.

Block-size workaround: ORT silently returns zeros for block_size > 256 (#28551). Both representations therefore expose block_size = 128 by replicating each native 512-element scale across 4 sub-blocks.

Other infrastructure changes

  • New flag tencent_q1_0_use_native_2bit (default False).
  • QuantizedLinear accepts bits ∈ {2, 4, 8} and a new zero_point_dtype argument. UINT8 (default) keeps the bit-packed integer form; float dtypes produce one un-packed value per block.
  • QuantizationConfig gains float_zero_point; TextModel passes config.dtype as the zp dtype when it is set.
  • GGUF → ArchitectureConfig derivation now sets rope_type from <arch>.rope.scaling.type (defaulting to "default" when rope.freq_base is present). Without this fix, models with rope.scaling.type = "none" were built with no RoPE at all.
  • For hunyuan-dense GGUFs whose rope.freq_base exceeds 1e6 (Tencent's pipeline bakes the dynamic-NTK exponent into a static value), the original HF config is restored: rope_type="dynamic", rope_theta=10000, alpha=1000. Otherwise long prompts diverge.
  • New GGUF arch mapping hunyuan-dense → hunyuan_v1_dense with attn_q_norm/attn_k_norm tensor name mappings.
  • _detect_quant_params returns block_size explicitly so the QuantizationConfig matches the actual per-type group size.

Verification

End-to-end on tencent/Hy-MT1.5-1.8B-2bit:

  • Per-tensor dequantization matches HF safetensors to bf16-rounding precision (max abs 2.3e-5).
  • Single-layer ORT matmul vs HF dequantized matmul: max abs 2.5e-3, mean 2.8e-4.
  • Short-prompt greedy generation with the chat template produces Hello, world!<EOS> token-for-token identical to HF.
  • Throughput with default flag: 0.13 s prefill, 32.8 tok/s decode on CPU EP.

Tests:

  • Split into TestTencentQ10DefaultInflated4Bit, TestTencentQ10NativeBits2, and TestTencentQ10Shared covering both flag values + shape/error invariants.
  • Full test suite (tests/build_graph_test.py src/): 2752 passed, 0 failures.

Out of scope / follow-ups

This adds two paths through the GGUF loader for GGML type 41:

1. **Mainline Q1_0** (1-bit binary, 18 bytes / 128 elements): inflated
   to MatMulNBits bits=2 with zero_point=1 so the {-d, +d} codebook
   becomes scale·(B - 1) for B ∈ {0, 2}. Verified against
   prism-ml/Bonsai-1.7B-Q1_0 reference checkpoints.

2. **Tencent custom Q1_0** (2-bit SEQ, 130 bytes / 512 elements):
   used by AngelSlim/Hy-MT1.5-1.8B-2bit-GGUF. The on-disk layout is
   per-row `[fp16 scale][2-bit codes packed LSB-first]` with the
   SEQ codebook {-1.5, -0.5, 0.5, 1.5} (stored scale = effective/2,
   equivalently codebook {-3, -1, +1, +3}·scale). Mainline llama.cpp
   refuses to load these files because every tensor after the first
   Q1_0 entry lands at the wrong offset — we bypass gguf-py's size
   calculation and read each tensor from its explicit per-tensor
   offset.

   The Tencent layout is inflated to MatMulNBits bits=4 with
   zero_point=3 (codes ∈ {0,2,4,6}) so dequant gives
   scale·(B - 3) ∈ {-3, -1, +1, +3}·scale exactly. We use bits=4
   instead of native bits=2 because the ORT CPU kernel currently
   only implements float zero_points for bits=4, and the SEQ offset
   1.5 cannot be expressed with an integer zp. The native 512-element
   block is replicated across four 128-element MatMulNBits sub-blocks
   (block_size=512 silently returns zeros on CPU EP).

Other infrastructure changes:

- QuantizedLinear now accepts bits ∈ {2, 4, 8}; zero_point packing
  generalized to `ceil(n_blocks * bits / 8)` bytes.
- GGUF→ArchitectureConfig conversion now derives `rope_type` from
  `<arch>.rope.scaling.type` (defaulting to "default" when
  `rope.freq_base` is present), fixing models whose graphs were
  previously built without RoPE.
- Added `hunyuan-dense` → `hunyuan_v1_dense` GGUF arch mapping and
  the corresponding attn_q_norm/attn_k_norm tensor name mappings.
- `_detect_quant_params` now returns block_size explicitly so the
  QuantizationConfig matches the actual per-type group size.

End-to-end verified on tencent/Hy-MT1.5-1.8B-2bit: greedy generation
produces `Hello, world!<EOS>` token-for-token identical to HF, and
per-tensor dequantization matches HF safetensors to bf16-rounding
precision (max abs 2.3e-5).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
@github-actions

github-actions Bot commented May 18, 2026 •

Copy link
Copy Markdown

Performance Comparison

Comparing 114a1bc → 9930a91

Model Metric Baseline Current Delta
bert (feature-extraction) model_size_bytes 359 KB 359 KB +0.0% ⚪
bert (feature-extraction) num_nodes 60 60 +0.0% ⚪
falcon model_size_bytes 364 KB 364 KB +0.0% ⚪
falcon num_nodes 66 66 +0.0% ⚪
gemma2 model_size_bytes 428 KB 428 KB +0.0% ⚪
gemma2 num_nodes 107 107 +0.0% ⚪
gpt2 model_size_bytes 388 KB 388 KB +0.0% ⚪
gpt2 num_nodes 53 53 +0.0% ⚪
llama model_size_bytes 425 KB 425 KB +0.0% ⚪
llama num_nodes 61 61 +0.0% ⚪
llama (static-cache) model_size_bytes 425 KB 425 KB +0.0% ⚪
llama (static-cache) num_nodes 58 58 +0.0% ⚪
mamba (ssm-text-generation) model_size_bytes 296 KB 296 KB +0.0% ⚪
mamba (ssm-text-generation) num_nodes 98 98 +0.0% ⚪
phi3 model_size_bytes 421 KB 421 KB +0.0% ⚪
phi3 num_nodes 59 59 +0.0% ⚪
phi3 (static-cache) model_size_bytes 421 KB 421 KB +0.0% ⚪
phi3 (static-cache) num_nodes 56 56 +0.0% ⚪
qwen2 model_size_bytes 425 KB 425 KB +0.0% ⚪
qwen2 num_nodes 61 61 +0.0% ⚪
qwen2 (static-cache) model_size_bytes 425 KB 425 KB +0.0% ⚪
qwen2 (static-cache) num_nodes 58 58 +0.0% ⚪
qwen3_5_moe (hybrid-text-generation) model_size_bytes 506 KB 506 KB +0.0% ⚪
qwen3_5_moe (hybrid-text-generation) num_nodes 275 275 +0.0% ⚪
qwen3_5_text (hybrid-text-generation) model_size_bytes 458 KB 458 KB +0.0% ⚪
qwen3_5_text (hybrid-text-generation) num_nodes 129 129 +0.0% ⚪
qwen3_5_vl (hybrid-qwen-vl) model_size_bytes 977 KB 977 KB +0.0% ⚪
qwen3_5_vl (hybrid-qwen-vl) num_nodes 413 413 +0.0% ⚪
t5 (seq2seq) model_size_bytes 836 KB 836 KB +0.0% ⚪
t5 (seq2seq) num_nodes 166 166 +0.0% ⚪
whisper (speech-to-text) model_size_bytes 1008 KB 1008 KB +0.0% ⚪
whisper (speech-to-text) num_nodes 128 128 +0.0% ⚪

No performance regressions.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds GGUF loader support for GGML_TYPE_Q1_0 (id 41) in two flavors: mainline llama.cpp 1-bit binary (used by prism-ml/Bonsai) and Tencent's custom 2-bit SEQ variant used by AngelSlim/Hy-MT1.5-1.8B-2bit-GGUF. Mainline Q1_0 is repacked into ORT MatMulNBits with bits=2, zp=1; Tencent Q1_0 is read via an offset-based parser and inflated into bits=4, zp=3 to work around ORT CPU kernel limitations on 2-bit float zero-points and 512-element block sizes. Several supporting infrastructure changes (generalized QuantizedLinear for bits=2, rope_type derivation from GGUF rope.scaling.type, hunyuan-dense arch mapping, and block_size propagation from _detect_quant_params) round out the change.

Changes:

  • Add Q1_0 (mainline 1-bit) repack path in _repacker.py plus extensive tests.
  • Add custom Tencent Q1_0 parser (_tencent_q1_0.py) that reads tensors from explicit GGUF offsets and produces 4-bit MatMulNBits arrays, integrated into _builder.py.
  • Generalize QuantizedLinear to support bits ∈ {2,4,8}, derive rope_type for GGUF, register hunyuan-dense, and surface block_size from quant detection.

Reviewed changes

Copilot reviewed 10 out of 10 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
src/mobius/integrations/gguf/_tensor_mapping.py Register hunyuan-dense/hunyuan_v1_dense mapping and add attn_q_norm/attn_k_norm Hunyuan extras.
src/mobius/integrations/gguf/_tencent_q1_0.py New Tencent Q1_0 layout detector and per-tensor parser producing MatMulNBits bits=4, zp=3 arrays.
src/mobius/integrations/gguf/_tencent_q1_0_test.py Unit tests for the Tencent parser: shape, scale replication, bit packing, and dequant round-trip.
src/mobius/integrations/gguf/_repacker.py Add mainline Q1_0 (1-bit) repack path → MatMulNBits 2-bit with shared zp=1.
src/mobius/integrations/gguf/_repacker_test.py New Q1_0 repack tests covering shape, packing, bit order, and round-trip dequant.
src/mobius/integrations/gguf/_config_mapping.py Derive rope_type from <arch>.rope.scaling.type and register Hunyuan model_type.
src/mobius/integrations/gguf/_builder.py Return block_size from quant detection, plumb Tencent Q1_0 path through quantized state-dict loading.
src/mobius/integrations/gguf/_builder_test.py Update _detect_quant_params signature usage and check block_size.
src/mobius/components/_quantized_linear.py Allow bits=2; generalize zero-point packing to ceil(n_blocks * bits / 8).
src/mobius/components/_quantized_linear_test.py Add tests for bits=2 shapes, zero-point packing, and forward attributes.

Comment thread src/mobius/integrations/gguf/_tencent_q1_0.py Outdated
Comment thread src/mobius/integrations/gguf/_tencent_q1_0.py Fixed
@github-actions

github-actions Bot commented May 18, 2026 •

Copy link
Copy Markdown

🏗️ Architecture Diff

Comparing 114a1bc → 9930a91

Model Sub-model Changes Status
bert (feature-extraction) model 0 ⚪
falcon model 0 ⚪
gemma2 model 0 ⚪
gemma4 (gemma4) decoder 0 ⚪
gemma4 (gemma4) embedding 0 ⚪
gemma4 (gemma4) vision_encoder 0 ⚪
gemma4_text model 0 ⚪
gpt2 model 0 ⚪
llama model 0 ⚪
llama (static-cache) model 0 ⚪
mamba (ssm-text-generation) model 0 ⚪
phi3 model 0 ⚪
phi3 (static-cache) model 0 ⚪
qwen model 0 ⚪
qwen (static-cache) model 0 ⚪
qwen2 model 0 ⚪
qwen2 (static-cache) model 0 ⚪
qwen2_moe model 0 ⚪
qwen2_moe (static-cache) model 0 ⚪
qwen3 model 0 ⚪
qwen3 (static-cache) model 0 ⚪
qwen3_5_moe (hybrid-text-generation) model 0 ⚪
qwen3_5_text (hybrid-text-generation) model 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) decoder 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) embedding 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) vision_encoder 0 ⚪
qwen3_moe model 0 ⚪
qwen3_moe (static-cache) model 0 ⚪
qwen3_next (hybrid-text-generation) model 0 ⚪
t5 (seq2seq) decoder 0 ⚪
t5 (seq2seq) encoder 0 ⚪
whisper (speech-to-text) decoder 0 ⚪
whisper (speech-to-text) encoder 0 ⚪

No architecture changes detected. ✅


Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed)

justinchuby and others added 2 commits May 18, 2026 15:39
Co-authored-by: Copilot Autofix powered by AI <223894421+github-code-quality[bot]@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchuby@users.noreply.github.com>
Earlier wording implied MatMulNBits doesn't support bits=2 at all.
Spec actually allows bits in {2..8}, and bits=2 with integer
zero_points works on the CPU EP. The two issues that actually force
the bits=4 / block_size=128 workarounds for SEQ are:

  1. The unpacked (float) zero_point CPU path is only implemented
     for bits=4 — needed because SEQ's offset 1.5 isn't an integer.
  2. bits=4, block_size=512 silently returns zeros on CPU EP.

Both have follow-up ORT issues. Updated module + builder docstrings
to reflect this precisely.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
Replace the bits=4 inflation workaround with the native 2-bit
representation. The on-disk weight bytes go through unchanged
(LSB-first 4-codes-per-byte matches MatMulNBits), and the SEQ
codebook { -1.5, -0.5, +0.5, +1.5 } * stored_scale is expressed as

    effective_scale = 2 * stored_scale
    zero_point      = 1.5
    dequant         = effective_scale * (code - 1.5)
                    = stored_scale * {-3, -1, +1, +3}[code]

This relies on the bits=2 float zero-point CPU fallback added in
microsoft/onnxruntime#28354 (merged 2026-05-13, shipping in ORT
1.27+). Single-layer parity verified: dequantized matmul against
the HF safetensors weight matches to max_abs=2.5e-3 (bf16 rounding).
Short-prompt greedy generation reproduces HF's translation
token-for-token through EOS.

Infrastructure changes:
  - QuantizedLinear gains a zero_point_dtype argument; UINT8 keeps
    the bit-packed integer form, float dtypes give one un-packed
    value per block.
  - QuantizationConfig gains float_zero_point; TextModel passes
    config.dtype as zero_point_dtype when it is set.
  - GGUF builder sets float_zero_point=True for Tencent Q1_0
    layouts (bits=2, block_size=128 because ORT silently zeros at
    block_size > 256 — see microsoft/onnxruntime#28551).
  - GGUF -> ArchitectureConfig restores HF dynamic-NTK RoPE
    (rope_theta=10000, alpha=1000) for hunyuan-dense files where
    Tencent baked a large freq_base (>1e6) into the metadata.
    Without this, long prompts diverge because the dynamic exponent
    isn't reapplied per position.

Test coverage updated to reflect the new tensor shapes:
bits=2, block_size=128, float32 zero_points = 1.5, scales =
2 * stored_scale.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
ORT 1.27's bits=2 + float-zp CPU kernel is a naive scalar fallback
("naive implementation, need to be optimized" per matmul_nbits.cc),
giving ~0.24 tok/s decode on this model vs ~32 tok/s for the bits=4
packed path -- ~130x slower on the same weights. Until an MLAS fused
kernel lands (tracked in microsoft/onnxruntime#28552), default to the
fast inflated representation and put the smaller native form behind
the MOBIUS_TENCENT_Q1_0_USE_NATIVE_2BIT flag.

Changes:
  - New flag tencent_q1_0_use_native_2bit (default False).
  - parse_tencent_q1_0_tensor splits into _read_tencent_blocks +
    _pack_inflated_4bit / _pack_native_2bit; dispatches on flag.
  - GGUF builder feeds float_zero_point=True only when the flag is on.
  - _detect_quant_params returns (bits=4, block_size=128) by default
    and (bits=2, block_size=128) when the flag is on.
  - Test classes split: TestTencentQ10DefaultInflated4Bit for the
    fast default, TestTencentQ10NativeBits2 for the flag-on path,
    TestTencentQ10Shared for shape/error tests.

End-to-end on tencent/Hy-MT1.5-1.8B-2bit with default flag:
prefill 0.13s, decode 32.8 tok/s on CPU EP.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
Comment thread src/mobius/integrations/gguf/_tencent_q1_0.py Fixed

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 13 out of 13 changed files in this pull request and generated no new comments.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
- Replace ambiguous unicode (-, x, /, etc.) with ASCII equivalents
  in changed-file docstrings and comments (RUFF/RUF002, RUF003).
- Rename uppercase N -> n in test locals (RUFF/N806).
- Reflow native-path test docstring summary (RUFF/D205).
- Apply ruff format auto-fixes across the PR diff.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
@justinchuby
justinchuby force-pushed the add-q1-0-gguf-support branch from c7f9343 to bce4293 Compare May 19, 2026 23:18
ORT GenAI's LLM model_type registry (onnxruntime-genai
src/models/model_type.h) does not include hunyuan_v1_dense, so the
genai_config.json mobius wrote for the Hy-MT1.5-1.8B-2bit export
failed to load. The registry does include 'decoder' as a generic
catch-all for any decoder-only causal LM that follows the standard
input/output shapes, which is exactly what hunyuan_v1_dense is.

This patch adds the mapping and bundles a runnable example
(examples/hy_mt1_5.py) that exercises the full pipeline:

  build (bf16 or Q1_0) -> ModelPackage.save ->
  write_ort_genai_config -> og.Model / og.Tokenizer / og.Generator

End-to-end verified on /tmp/hymt-q1-genai (Q1_0 quantized export,
CPU EP, ORT GenAI 0.13.2): session loads, generator runs,
generate_next_token() returns logical tokens. The example also
patches past_present_share_buffer to False post-hoc because mobius
currently emits share_buffer=True for the CPU EP even though our
exports use a dynamic-shape KV cache; that mismatch is documented
inline as a separate follow-up.

A known downstream caveat is documented in the example header:
ort-extensions' BPE tokenizer does not round-trip the upstream
Hy-MT1.5 vocab for CJK content (e.g. '你好' tokenizes to an empty
sequence), so the model is fed an English-only prompt and replies
'Please provide the text you want to be translated.' The decoder
loads and generates correctly; the tokenizer parity is independent
of this PR.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
@justinchuby
justinchuby force-pushed the add-q1-0-gguf-support branch from bce4293 to cb26b68 Compare May 19, 2026 23:18
The example originally used og.Tokenizer, but the Hy-MT1.5 vocab's
custom regex pre-tokenizer is not currently round-tripped by
ort-extensions (e.g. 'Hello, world!' tokenizes to a single space
token, '你好' tokenizes to []), so the model never saw the actual
prompt content. By default we now tokenize with HF and feed raw
token IDs to og.Generator -- still exercises the full ORT GenAI
inference path. Pass --use-ort-tokenizer to reproduce the broken
path for debugging.

Verified working end-to-end on the Q1_0 quantized export, CPU EP:

  'Translate to Spanish: The cat is sleeping on the chair.'
    -> 'El gato está durmiendo en la silla.'

  'Translate to French: Knowledge is power.'
    -> 'La connaissance est puissance.'

  'Translate the following Chinese text to English: 你好,世界!'
    -> 'Hello, world!'

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
@justinchuby
justinchuby force-pushed the add-q1-0-gguf-support branch from 16162c5 to fb59741 Compare May 19, 2026 23:25
…-gguf

ORT GenAI's past_present_share_buffer mode requires the decoder graph
to update the KV cache in place (com.microsoft.GroupQueryAttention).
The standard ONNX opset 23 Attention op concatenates past_key + new_K
and returns a dynamic-shape present_key, which is incompatible with
the pre-allocated shared buffer ORT GenAI passes in. The previous
config wrote past_present_share_buffer=True purely from the target EP
flag, which meant builds that didn't apply the GQA rewrite (anything
other than --ep cpu / cuda / dml / etc.) silently shipped a broken
config and crashed at first inference with 'inconsistent
total_sequence_length'.

This change:

  * Adds supports_in_place_kv_cache to GenaiConfigGenerator and
    _default_search_params; when set it overrides the EP flag.
  * In auto_export._write_genai_config, introspects the decoder graph
    for GroupQueryAttention nodes and passes the result through.
    Standard-Attention graphs now correctly get share_buffer=False
    regardless of the target EP.
  * Adds --ep to the build-gguf CLI and threads execution_provider
    through build_from_gguf so users can opt into the GQA rewrite for
    GGUF imports the same way they can for HF imports.
  * Updates examples/hy_mt1_5.py to build the Q1_0 variant with
    --ep cpu so the example exercises the full share-buffer path
    end-to-end and removes the post-hoc _patch_genai_config workaround.

Verified end-to-end on tencent/Hy-MT1.5-1.8B-2bit:
  - 32 GroupQueryAttention nodes after build
  - genai_config.search.past_present_share_buffer = true
  - 'Translate to Spanish: The cat is sleeping on the chair.'
    -> 'El gato está durmiendo en la silla.'

Full suite: 2753 passed, 0 failures.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Justin Chu <justinchu@microsoft.com>
@justinchuby
justinchuby merged commit 42573f4 into main May 20, 2026
20 of 23 checks passed
@justinchuby
justinchuby deleted the add-q1-0-gguf-support branch May 20, 2026 00:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants