Skip to content

Add experimental INT3 QMoE checkpoint export - #764

Open
titaiwangms wants to merge 2 commits into
mainfrom
draft/int3-qmoe-export
Open

titaiwangms wants to merge 2 commits into
mainfrom
draft/int3-qmoe-export

Conversation

@titaiwangms

Copy link
Copy Markdown
Contributor

Summary

Add a default-off, experimental exporter for Olive fused INT3 MoE checkpoints.
This is an export-only draft, not a claim that released ORT can execute INT3
QMoE models. It follows the proposed raw-weight direction in
microsoft/onnxruntime#32657 and consumes the native Torch INT3 checkpoint format
introduced by microsoft/Olive#2716.

Changes

  • Gate export behind MOBIUS_EXPERIMENTAL_INT3_QMOE_EXPORT=1 or
    override_flags(experimental_int3_qmoe_export=True).
  • Support FC1/FC2 pairs (3,4), (3,3), and (4,3) in the generic MoE adapter,
    including Qwen3-MoE, with exact Olive expert overrides and an INT4 model-wide
    fallback.
  • Preserve row-local, LSB-first packed unsigned INT3 codes with symmetric
    zero point 4. Interleave FC1 gate/up bytes and scales during preprocessing;
    preserve FC2's K-last layout without requantization.
  • Validate common block sizes 32/64/128, dimension alignment, FP16/BF16
    activation/scale agreement, finite nonnegative scales, and optional explicit
    INT3 zero points. Preserve existing INT4 zero-point semantics.
  • Emit explicit FC bit-width attributes and draft-format/runtime-warning
    metadata, without assigning an unapproved schema version.
  • Add independent packing/dequantization oracles, negative validation cases,
    local-checkpoint public-build coverage, and external-data save/reload tests.

Validation

Run from the isolated draft worktree:

PYTHONPATH="$PWD/src" HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 \
  python -m pytest src/mobius/models/int3_qmoe_test.py \
  src/mobius/models/moe_test.py src/mobius/_weight_utils_test.py \
  src/mobius/components/_moe_test.py -q --tb=short

Passed: 312 tests, no skips, including 57 new INT3 tests.

lintrunner f --output oneline README.md src/mobius/_flags.py \
  src/mobius/_weight_utils.py src/mobius/components/_moe.py \
  src/mobius/models/moe.py src/mobius/models/int3_qmoe_test.py

Passed.

git diff --check

Passed.

Draft boundaries

  • No INT3 ORT kernel execution, checkpoint-to-runtime logits parity, real-model
    quality, or performance qualification was run.
  • Dense INT3 MatMulNBits, quantized embeddings, (3,8)/(8,3), INT2/INT3 mixed
    pairs, and other checkpoint layouts are not enabled.
  • Final QMoE schema/version agreement, complete quantization provenance, and
    runtime qualification remain follow-up gates. This does not complete the
    end-to-end plan's P5 acceptance gate.
  • INT2 runtime/CI qualification is separate and is not included in this PR.

Gate raw Olive INT3 fused expert export behind an experimental flag. Validate packing, scales and source profiles, and cover public local-checkpoint export and external-data round trips. Runtime schema and kernel qualification remain pending.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Signed-off-by: titaiwang <titaiwang@microsoft.com>
@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Performance Comparison

Comparing 88f7d045 → 2101f88a

Model Metric Baseline Current Delta
bert (feature-extraction) model_size_bytes 359 KB 359 KB +0.0% ⚪
bert (feature-extraction) num_nodes 68 68 +0.0% ⚪
falcon model_size_bytes 364 KB 364 KB +0.0% ⚪
falcon num_nodes 66 66 +0.0% ⚪
gemma2 model_size_bytes 428 KB 428 KB +0.0% ⚪
gemma2 num_nodes 105 105 +0.0% ⚪
gpt2 model_size_bytes 324 KB 324 KB +0.0% ⚪
gpt2 num_nodes 54 54 +0.0% ⚪
llama model_size_bytes 425 KB 425 KB +0.0% ⚪
llama num_nodes 60 60 +0.0% ⚪
llama (static-cache) model_size_bytes 425 KB 425 KB +0.0% ⚪
llama (static-cache) num_nodes 56 56 +0.0% ⚪
mamba (ssm-text-generation) model_size_bytes 296 KB 296 KB +0.0% ⚪
mamba (ssm-text-generation) num_nodes 94 94 +0.0% ⚪
phi3 model_size_bytes 421 KB 421 KB +0.0% ⚪
phi3 num_nodes 58 58 +0.0% ⚪
phi3 (static-cache) model_size_bytes 421 KB 421 KB +0.0% ⚪
phi3 (static-cache) num_nodes 54 54 +0.0% ⚪
qwen2 model_size_bytes 425 KB 425 KB +0.0% ⚪
qwen2 num_nodes 60 60 +0.0% ⚪
qwen2 (static-cache) model_size_bytes 425 KB 425 KB +0.0% ⚪
qwen2 (static-cache) num_nodes 56 56 +0.0% ⚪
qwen3_5_moe (hybrid-text-generation) model_size_bytes 506 KB 506 KB +0.0% ⚪
qwen3_5_moe (hybrid-text-generation) num_nodes 265 265 +0.0% ⚪
qwen3_5_text (hybrid-text-generation) model_size_bytes 458 KB 458 KB +0.0% ⚪
qwen3_5_text (hybrid-text-generation) num_nodes 127 127 +0.0% ⚪
qwen3_5_vl (hybrid-qwen-vl) model_size_bytes 977 KB 977 KB +0.0% ⚪
qwen3_5_vl (hybrid-qwen-vl) num_nodes 450 450 +0.0% ⚪
t5 (seq2seq) model_size_bytes 836 KB 836 KB +0.0% ⚪
t5 (seq2seq) num_nodes 176 176 +0.0% ⚪
whisper (speech-to-text) model_size_bytes 1008 KB 1008 KB +0.0% ⚪
whisper (speech-to-text) num_nodes 128 128 +0.0% ⚪

No performance regressions.

@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

🏗️ Architecture Diff

Comparing 88f7d045 → 2101f88a

Model Sub-model Changes Status
bert (feature-extraction) model 0 ⚪
falcon model 0 ⚪
gemma2 model 0 ⚪
gemma4 (gemma4) decoder 0 ⚪
gemma4 (gemma4) embedding 0 ⚪
gemma4 (gemma4) vision_encoder 0 ⚪
gemma4_text model 0 ⚪
gpt2 model 0 ⚪
llama model 0 ⚪
llama (static-cache) model 0 ⚪
mamba (ssm-text-generation) model 0 ⚪
phi3 model 0 ⚪
phi3 (static-cache) model 0 ⚪
qwen model 0 ⚪
qwen (static-cache) model 0 ⚪
qwen2 model 0 ⚪
qwen2 (static-cache) model 0 ⚪
qwen2_moe model 0 ⚪
qwen2_moe (static-cache) model 0 ⚪
qwen3 model 0 ⚪
qwen3 (static-cache) model 0 ⚪
qwen3_5_moe (hybrid-text-generation) model 0 ⚪
qwen3_5_text (hybrid-text-generation) model 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) decoder 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) embedding 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) vision_encoder 0 ⚪
qwen3_moe model 0 ⚪
qwen3_moe (static-cache) model 0 ⚪
qwen3_next (hybrid-text-generation) model 0 ⚪
t5 (seq2seq) decoder 0 ⚪
t5 (seq2seq) encoder 0 ⚪
whisper (speech-to-text) decoder 0 ⚪
whisper (speech-to-text) encoder 0 ⚪

No architecture changes detected. ✅


Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed)

@titaiwangms
titaiwangms marked this pull request as ready for review October 9, 2026 18:46
@titaiwangms
titaiwangms requested review from a team and a balanced review from Copilot October 9, 2026 18:46

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

The documented programmatic opt-in relies on a private, test-oriented API rather than a supported public export.

2 open findings
What changed in this PR

Adds an opt-in experimental exporter for Olive INT3 QMoE checkpoints while explicitly deferring runtime support.

Changes:

  • Adds INT3 layout validation, packing, and QMoE attributes.
  • Introduces an experimental feature flag and documentation.
  • Adds comprehensive offline export and round-trip tests.
File Description
README.md Documents experimental usage and limitations.
src/​mobius/​_flags.py Adds the INT3 export feature flag.
src/​mobius/​_weight_utils.py Validates and preprocesses INT3 checkpoint data.
src/​mobius/​components/​_moe.py Emits draft INT3 QMoE graphs and metadata.
src/​mobius/​models/​moe.py Passes activation dtype into preprocessing.
src/​mobius/​models/​int3_qmoe_test.py Tests packing, validation, builds, and persistence.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread README.md

```python
from mobius import build
from mobius._flags import override_flags
Comment on lines +290 to +291
"INT3 QMoE export requires experimental_int3_qmoe_export=True; "
"the proposed INT3 schema and kernels are not released runtime support."

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants