Skip to content

Add Clef-Flash joint schema decision model support - #760

Draft
gramalingam wants to merge 3 commits into
mainfrom
gramalingam-clef-flash-support
Draft

gramalingam wants to merge 3 commits into
mainfrom
gramalingam-clef-flash-support

Conversation

@gramalingam

Copy link
Copy Markdown
Collaborator

Summary

Add support for Cloudflare/clef-flash as a full-record structured decision model rather than an autoregressive language model.

  • Export vision encoder, independent image/video embedding mixer, hidden-state backbone, and joint schema decision head.
  • Detect the official checkpoint and local head sidecars before ordinary Qwen3.5 dispatch; pin the default release to fde727a287004204b7518dcc983fe64379776712 and forward revisions across assets.
  • Stream checkpoint weights, including splitting the head's fused QKV tensors and using the output LM embedding table for lexical features.
  • Reproduce publisher schema encoding and labeled option probabilities; reject incompatible autoregressive options and emit only advisory onnx-genai contracts.
  • Match nested rotary settings and DeltaNet precision, accumulate schema pooling in float32, and cast processor pixels to vision weight dtype.
  • Add documentation and a text-only ONNX Runtime inference example.

Validation

  • Final Clef and shared Qwen vision selection: 69 passed, 17 skipped. Covers float32/float16 streamed tiny-checkpoint parity for text, image, video, and mixed-media records; pinned publisher head/schema/processor comparisons; actual published-config graph construction; and loader, revision, CLI, and metadata checks.
  • Registry, Transformers builder/config resolver, runtime exporters, and core graph coverage: 1,465 passed, 33 skipped.
  • CLI regressions: 79 passed; existing Qwen3.5 text synthetic parity: 2 passed.
  • Strict typing passes for the four new source modules. Formatting and available Ruff lint rules pass; the installed Ruff does not recognize repository selector RUF105, so equivalent rules were run without that selector.

Limitations and waivers

  • Full real-weight 9B inference and decision goldens have not been established. Numerical end-to-end evidence uses reduced-size, randomly initialized backbones and the pinned publisher's head implementation.
  • BF16 numerical execution requires CUDA kernels and is skipped on CPU. Full-pipeline BF16, CUDA/DML, and Olive quantization are unverified.
  • Generic causal-LM L4/L5 generation tests do not implement this decision contract; dedicated Clef coverage is used instead. No decision-specific golden fixture is claimed.
  • Padding/batching, DeepStack, quantized source checkpoints, and autoregressive runtime configuration are unsupported. ORT GenAI config is explicitly rejected.

Keeping this PR in draft pending full-size checkpoint and accelerator validation.

Export the pinned Cloudflare checkpoint as vision, embedding, full-record backbone, and joint decision-head components. Preserve publisher schema spans and option probabilities, stream backbone and fused-QKV head weights, and reject autoregressive runtime configurations.

Match nested Qwen3.5 rotary settings and DeltaNet precision, pool long schema spans in float32, and cast processor pixels to the vision weight dtype. Add publisher parity, streamed multimodal pipeline, revision, loader, CLI, and metadata coverage plus documentation and a text inference example. Full-size real-checkpoint and accelerator inference remain explicitly unverified.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ganesan Ramalingam <grama@microsoft.com>
@CLAassistant

Copy link
Copy Markdown

CLA assistant check
Thank you for your submission! We really appreciate it. Like many open source projects, we ask that you sign our Contributor License Agreement before we can accept your contribution.
You have signed the CLA already but the status is still pending? Let us recheck it.

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Performance Comparison

Comparing dc179846 → ae87fa0c

Model Metric Baseline Current Delta
bert (feature-extraction) model_size_bytes 359 KB 359 KB +0.0% ⚪
bert (feature-extraction) num_nodes 68 68 +0.0% ⚪
falcon model_size_bytes 364 KB 364 KB +0.0% ⚪
falcon num_nodes 66 66 +0.0% ⚪
gemma2 model_size_bytes 428 KB 428 KB +0.0% ⚪
gemma2 num_nodes 105 105 +0.0% ⚪
gpt2 model_size_bytes 324 KB 324 KB +0.0% ⚪
gpt2 num_nodes 54 54 +0.0% ⚪
llama model_size_bytes 425 KB 425 KB +0.0% ⚪
llama num_nodes 60 60 +0.0% ⚪
llama (static-cache) model_size_bytes 425 KB 425 KB +0.0% ⚪
llama (static-cache) num_nodes 56 56 +0.0% ⚪
mamba (ssm-text-generation) model_size_bytes 296 KB 296 KB +0.0% ⚪
mamba (ssm-text-generation) num_nodes 94 94 +0.0% ⚪
phi3 model_size_bytes 421 KB 421 KB +0.0% ⚪
phi3 num_nodes 58 58 +0.0% ⚪
phi3 (static-cache) model_size_bytes 421 KB 421 KB +0.0% ⚪
phi3 (static-cache) num_nodes 54 54 +0.0% ⚪
qwen2 model_size_bytes 425 KB 425 KB +0.0% ⚪
qwen2 num_nodes 60 60 +0.0% ⚪
qwen2 (static-cache) model_size_bytes 425 KB 425 KB +0.0% ⚪
qwen2 (static-cache) num_nodes 56 56 +0.0% ⚪
qwen3_5_moe (hybrid-text-generation) model_size_bytes 506 KB 506 KB +0.0% ⚪
qwen3_5_moe (hybrid-text-generation) num_nodes 265 265 +0.0% ⚪
qwen3_5_text (hybrid-text-generation) model_size_bytes 458 KB 458 KB +0.0% ⚪
qwen3_5_text (hybrid-text-generation) num_nodes 127 127 +0.0% ⚪
qwen3_5_vl (hybrid-qwen-vl) model_size_bytes 977 KB 977 KB +0.0% ⚪
qwen3_5_vl (hybrid-qwen-vl) num_nodes 450 451 +0.2% ⚪
t5 (seq2seq) model_size_bytes 836 KB 836 KB +0.0% ⚪
t5 (seq2seq) num_nodes 176 176 +0.0% ⚪
whisper (speech-to-text) model_size_bytes 1008 KB 1008 KB +0.0% ⚪
whisper (speech-to-text) num_nodes 128 128 +0.0% ⚪

No performance regressions.

Comment thread src/mobius/models/clef_test.py Fixed
Comment thread src/mobius/models/clef_test.py Fixed
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

🏗️ Architecture Diff

Comparing dc179846 → ae87fa0c

Model Sub-model Changes Status
bert (feature-extraction) model 0 ⚪
falcon model 0 ⚪
gemma2 model 0 ⚪
gemma4 (gemma4) decoder 0 ⚪
gemma4 (gemma4) embedding 0 ⚪
gemma4 (gemma4) vision_encoder 0 ⚪
gemma4_text model 0 ⚪
gpt2 model 0 ⚪
llama model 0 ⚪
llama (static-cache) model 0 ⚪
mamba (ssm-text-generation) model 0 ⚪
phi3 model 0 ⚪
phi3 (static-cache) model 0 ⚪
qwen model 0 ⚪
qwen (static-cache) model 0 ⚪
qwen2 model 0 ⚪
qwen2 (static-cache) model 0 ⚪
qwen2_moe model 0 ⚪
qwen2_moe (static-cache) model 0 ⚪
qwen3 model 0 ⚪
qwen3 (static-cache) model 0 ⚪
qwen3_5_moe (hybrid-text-generation) model 0 ⚪
qwen3_5_text (hybrid-text-generation) model 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) decoder 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) embedding 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) vision_encoder 35 🟡
qwen3_moe model 0 ⚪
qwen3_moe (static-cache) model 0 ⚪
qwen3_next (hybrid-text-generation) model 0 ⚪
t5 (seq2seq) decoder 0 ⚪
t5 (seq2seq) encoder 0 ⚪
whisper (speech-to-text) decoder 0 ⚪
whisper (speech-to-text) encoder 0 ⚪
qwen3_5_vl (hybrid-qwen-vl) / vision_encoder — 35 change(s)

Op summary: 251 → 252 nodes

--- base
+++ head
@@ -1,3 +1,4 @@
+Cast
 Reshape
 Conv
 Reshape

Added nodes:

  • + Cast

Modified attributes:

  • node[93] Constant: value_int: 0 → None, value_ints: None → [1, 0]
  • node[127] Constant: value_int: 1 → 0
  • node[202] Constant: value_int: 1 → 0

Connectivity changes:

  • node[10] Mul: input_ids [61, 66] → [64, 66]
  • node[21] Unsqueeze: input_ids [67, 8] → [77, 7]
  • node[39] Mul: input_ids [90, 95] → [93, 95]
  • node[50] Unsqueeze: input_ids [96, 8] → [106, 7]
  • node[60] Gather: input_ids [115, 8] → [116, 7]
  • node[64] Gather: input_ids [13, 118] → [12, 119]
  • node[65] Gather: input_ids [12, 119] → [13, 119]
  • node[66] Gather: input_ids [13, 119] → [12, 120]
  • node[81] Unsqueeze: input_ids [127, 8] → [137, 7]
  • node[103] CastLike: input_ids [125, 159] → [125, 160]
  • node[112] Mul: input_ids [163, 166] → [165, 166]
  • node[119] Mul: input_ids [163, 168] → [165, 168]
  • node[123] Reshape: input_ids [182, 20] → [176, 20]
  • node[130] Unsqueeze: input_ids [151, 7] → [190, 8]
  • node[137] Unsqueeze: input_ids [196, 7] → [197, 8]
  • node[140] CastLike: input_ids [22, 183] → [21, 184]
  • node[143] Unsqueeze: input_ids [183, 7] → [203, 23]
  • node[145] Unsqueeze: input_ids [158, 7] → [185, 7]
  • node[151] Add: input_ids [88, 211] → [211, 25]
  • node[160] Add: input_ids [212, 220] → [220, 31]
  • node[178] CastLike: input_ids [125, 238] → [125, 239]
  • node[187] Mul: input_ids [242, 245] → [244, 245]
  • node[194] Mul: input_ids [242, 247] → [244, 247]
  • node[198] Reshape: input_ids [261, 20] → [255, 20]
  • node[205] Unsqueeze: input_ids [151, 7] → [269, 8]
  • node[212] Unsqueeze: input_ids [275, 7] → [276, 8]
  • node[215] CastLike: input_ids [22, 262] → [21, 263]
  • node[218] Unsqueeze: input_ids [262, 7] → [282, 23]
  • node[220] Unsqueeze: input_ids [237, 7] → [264, 7]
  • node[226] Add: input_ids [221, 290] → [290, 44]
  • node[235] Add: input_ids [291, 299] → [299, 50]

Legend: ⚪ No change · 🔵 Minor (attrs/inits) · 🟡 Moderate (nodes added/removed) · 🔴 Major (interface changed)

Record the dedicated four-component Clef L1-L3 tests in the coverage audit exemption list. The generic causal-LM/config harness cannot represent schema span inputs or load the separate head sidecar; keep the real-checkpoint golden limitation explicit instead of claiming generic coverage.

Sort Qwen vision imports using the CI-pinned Ruff 0.16.2 rules. Validate the coverage audit and complete Clef suite, plus lint and formatting across all PR Python changes.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ganesan Ramalingam <grama@microsoft.com>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Cross-dtype fused-QKV serialization can fail, and registry coverage metadata is incomplete.

3 open findings
What changed in this PR

Adds Clef-Flash as a four-component multimodal structured-decision pipeline.

Changes:

  • Adds backbone, vision, embedding, decision-head export and weight streaming.
  • Adds host-side schema encoding and runtime contracts.
  • Adds tests, documentation, and an ONNX Runtime example.
File Description
tests/​build_graph/​_support.py Exempts specialized Clef I/O.
src/​mobius/​tasks/​_vision_language_3model.py Casts processor pixels to model dtype.
src/​mobius/​tasks/​_clef.py Builds the four Clef components.
src/​mobius/​tasks/​__init__.py Registers the Clef task.
src/​mobius/​models/​qwen35.py Broadens embedding model typing.
src/​mobius/​models/​clef.py Implements Clef configuration, head, and model.
src/​mobius/​models/​clef_test.py Adds graph, parity, loader, and runtime tests.
src/​mobius/​models/​__init__.py Exports ClefFlashModel.
src/​mobius/​integrations/​transformers/​_clef.py Detects and streams Clef checkpoints.
src/​mobius/​integrations/​transformers/​_builder.py Dispatches Clef before Qwen3.5.
src/​mobius/​integrations/​ort_genai/​auto_export.py Rejects unsupported ORT GenAI export.
src/​mobius/​integrations/​onnx_genai/​auto_export.py Emits advisory component metadata.
src/​mobius/​integrations/​clef.py Encodes schemas and restores labels.
src/​mobius/​_registry.py Registers the Clef architecture.
src/​mobius/​__main__.py Routes Clef runtime metadata correctly.
README.md Lists structured-decision support.
examples/​clef_flash.py Demonstrates text-only inference.
docs/​model-catalog.md Adds Clef to the catalog.
docs/​index.md Includes the Clef guide.
docs/​clef-flash.md Documents export, inference, and limitations.

🧠 Review effort: Balanced


💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/mobius/_registry.py
Comment thread src/mobius/integrations/transformers/_clef.py Outdated
Comment thread src/mobius/models/clef_test.py
Cast each fused attention QKV split to the target initializer dtype before the lazy streaming loader materializes it. Cover all nine float32/float16/bfloat16 source-target combinations for weights and biases through save/reload.

Run multimodal pipeline parity on reloaded exported packages rather than original lazy bindings, and consistently import the public build function in Clef tests.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Signed-off-by: Ganesan Ramalingam <grama@microsoft.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants