Skip to content

Finegrained kernels: unified quantization kernels api - #48058

Open
IlyasMoutawwakil wants to merge 44 commits into
mainfrom
finegrained-unified
Open

IlyasMoutawwakil wants to merge 44 commits into
mainfrom
finegrained-unified

Conversation

@IlyasMoutawwakil

@IlyasMoutawwakil IlyasMoutawwakil commented Aug 18, 2026 •

Copy link
Copy Markdown
Member

CPU CI GPU run-slow

What does this PR do?

Fixes # (issue)

Code Agent Policy

The Transformers repo is currently being overwhelmed by a large number of PRs and issue comments written by
code agents. These often are low-quality, or fix extremely minor issues that occur rarely or never in practice.
As a result, we're instituting a rule that first-time contributors should not use code agents to submit PRs or issues.
We'd also ask autonomous "OpenClaw"-like agents not to open any PRs or issues.

Issues/PRs from first-time contributors that violate this rule will probably just be closed without review, and we
might block you, especially if you open more than one or appear to be deliberately ignoring this. We especially do not
want new contributors to jump in on random issues to contribute an agent-written fix. This creates lots of noise
for reviewers and other users and will almost certainly get you blocked.

For more information, please read CONTRIBUTING.md.

  • (First-time contributors only): I confirm that this PR description and code is not written by an LLM or code agent

Before submitting

  • This PR fixes a typo or improves the docs (you can dismiss the other checks if that's the case).
  • Did you read the contributor guideline and the
    Pull Request checks?
  • Was this discussed/approved via a Github issue or the forum? Please add a link
    to it if that's the case.
  • Did you make sure to update the documentation with your changes according to the guidelines?
  • Did you write any new necessary tests?

Who can review?

Anyone in the community is free to review the PR once the tests have passed. Feel free to tag
members/contributors who may be interested in your PR.

The kernels now store gate|up row-interleaved (gate j at row 2j, up j at 2j+1)
rather than as two N-apart halves, so the integration follows:

  - GPT-OSS already ships this order on disk, so its de-interleave converters go
    away entirely: _deinterleave_gate_up_rows, its use in FineGrainedMxfp4Deserialize,
    and the whole FineGrainedGateUpBiasDeinterleave op plus its registration.
  - Checkpoints that ship gate_up stacked are interleaved once post-load by
    interleave_gate_up_after_loading, ordered BEFORE swizzle_scales_after_loading
    (the swizzle packs whatever row order it finds) and skipped for mxfp4. This
    cannot live in conversion_mapping's MergeModulelist+Concatenate chain — that is
    shared with non-finegrained MoE models.
  - _apply_gate de-interleaves with a stride-2 split, mirroring the fused epilogue's
    split_gate_up; the two must agree or fused and unfused stop being comparable.
  - swizzle_scales_after_loading drops gate= and gates on the doubled extent
    (n_rows % 128, i.e. N % 64), so GPT-OSS N=2880 now pre-swizzles instead of
    falling back to an affine scale read.

Adds tests/kernels/test_finegrained.py (24 tests), covering the post-load interleave,
that mxfp4 is left untouched, dispatch/swizzle gating, and the MergeModulelist path.

tests/kernels/test_finegrained.py: 24 passed.
The kernels store gate|up row-interleaved, but ``transform_weights_for_mega_moe``
does its OWN gate/up interleave and so takes the stacked form. Interleaving those
modules at load and undoing it at the boundary is a pointless round trip — and
interleaving twice is not the identity, it is a different permutation, i.e. silent
garbage rather than a crash.

So modules bound for megamoe are simply never interleaved: the post-load pass skips
them and records what the module actually holds in ``_gate_up_interleaved`` (the
bytes cannot be inspected to tell, and ``set_experts_implementation`` can switch
backends after load). ``setup_megamoe_weights`` reads that flag and raises if it is
ever handed an interleaved module, rather than re-permuting on an assumption.

The other DeepGEMM expert paths need no change: their grouped GEMM is layout-agnostic
and the gate split happens in ``_apply_gate``, which already de-interleaves stride-2.

Verified: triton backend interleaves and flags; megamoe backend is left untouched with
no round trip. The raise itself is unverified here — DeepGEMM's JIT needs a CUDA
toolkit >= 12.9 and CUDA_HOME is unset on this box, so no megamoe path executes.
@HuggingFaceDocBuilderDev

Copy link
Copy Markdown

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

IlyasMoutawwakil and others added 9 commits August 18, 2026 11:38
# Conflicts:
#	src/transformers/integrations/hub_kernels.py
#	src/transformers/quantizers/auto.py
The memoized `(epilogue, gate_up_quantization, down_quantization)` triple bought
nothing where it matters. Constructing the three dataclasses measured ~1.4us per
layer (~85us per step at 61 layers), and only in EAGER decode — under cudagraphs
the host path does not run on replay, which is the mode this integration deploys
in. Against that it held hidden mutable state on the module, and cached a gate_up
quantization carrying an `output_recipe` that silently mismatches any caller
asking for the unfused form.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The kernels now take a per-expert bias as a `bias=` operand, so a biased model
no longer falls back to the unfused two-GEMM form just for having a bias.
`_kernel_epilogue` stops bailing on `has_bias`, and both host-side adds go away:
`_apply_unfused_gate_up` keeps only the activation, `_finish_down` only the
routing-weighted reduce. That also drops the `torch.sort` the grouped path paid
to index the gate_up bias by expert-sorted row -- the kernel indexes by the tile's
own expert id instead.

An unsupported act_fn no longer costs the bias its fusion either: the bias is an
operand beside the epilogue rather than part of it, so it fuses whether or not
the GLU does, and `epilogue is None` goes back to meaning exactly one thing.

The bias rides the same output-row axis as the weight and scale grid, so
`interleave_gate_up_after_loading` already delivers it in the kernels'
interleaved order -- no reordering here.

Tests assert the operand reaches both GEMMs, plus a new case pinning that an
unfusable activation still fuses its bias.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Conflict resolutions:
- quantizers/auto.py: keep the finegrained claims on fp8/nvfp4/modelopt
  (pre-quantized + MoE experts; NVFP4HfQuantizer only quantizes on the fly),
  take main's new gguf entry.
- integrations/__init__.py: union — the finegrained module plus main's
  FP8Embedding exports.
- integrations/deepgemm.py: ours (to_local unwrap + the stacked-gate guard);
  re-add the to_local import main's shim conversion dropped.
- integrations/tensor_parallel.py: main's shim, with to_local restored on it —
  the kernel integrations still import it from this path.
- distributed/sharding_utils.py: re-apply the 0-dim scalar guard (ModelOpt
  per-projection weight_scale_2/input_scale) inside the rewritten
  DtensorShardOperation.shard_tensor: replicated on the dense path,
  expert-ownership-filtered on the MoE path.
- tests/tensor_parallel: main's rewrite, with the 0-dim sharding test re-added
  against the new engine (passes).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e kernels' MoE chain

The experts forwards are adapters over the kernels' `moe_fused_batched/grouped`: `_moe_operands`
hands them the module's tensors and the activation — a `get_supported_act_fns()` name the gate_up
epilogue fuses, else the module's own GLU as a callable the kernels run on the host, so any
activation works without a kernel release. The kernel bundle is loaded by symbol name
(`matmul_2d/batched/grouped`, the two forwards, swizzle/unswizzle, `get_supported_act_fns`).

Every layout difference between a checkpoint and what the modules hold is a `ConversionOps` with a
reverse, attached by the quantizer, so `save_pretrained` restores the checkpoint bitwise:
- `FineGrainedInterleaveGateUp`: stacked [gate; up] rows -> the kernels' [g0, u0, ...] order,
  skipped when the model declares `is_concatenated=False` (GPT-OSS; the flag travels through
  `use_experts_implementation`, which stamps its defaults after `__init__`) or the backend packs
  gate|up itself (DeepGEMM Mega MoE);
- `FineGrainedScaleContainer`: a scale into the dtype its module holds — the same bytes for a
  uint8 UE8M0 container (MiniMax), an exact cast for float32 values (dsv4-flash-base) — and the
  module records the shipped container so the reverse restores it;
- `FineGrainedSwizzleScales`: the SWIZZLE_32_4_4 artifact for modules that hold 5-D scales
  (`FineGrainedExperts.__init__` allocates that shape for every non-DeepGEMM backend on SM100
  with a group-scaled format and activation quant; whole 128-row/4-col blocks only);
- `FineGrainedPackedBlocks` regroups GPT-OSS `{proj}_blocks` into the packed weight, so the
  blocks/scales format is two one-to-one converters and round-trips too.
Catch-all converters cover keys that already arrive under the fused names and dense scale keys.
The post-load hooks, `_save_to_state_dict` override and post-load dtype cast are gone.

Also: the eager per-expert loop reads a swizzled stack's `[e:e+1]` slice and passes bias, NVFP4
global and activation format (it dropped all three); expert biases take the model dtype; the
`mxfp4` format pins UE8M0 scales; `FineGrainedEmbedding` (FP8 table, per-tensor scale) for
`modules_to_convert`; DeepGEMM backends refuse modules holding swizzled scales or interleaved rows
before loading the kernel; `disable_deepgemm_on_multi_device` is public.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…heckpoint precision for quantized params

`FineGrainedQuantize` reads the module's weight format: block-FP8 in torch (per-block E4M3, fp32 or
power-of-two UE8M0 inverse scales, from the module's block and held scale dtype), MXFP8 / MXFP4 /
NVFP4 through the kernels' row-wise quantizers in one launch per tensor (NVFP4 normalized by the
per-tensor / per-expert global `amax / (6 * 448)`), and emits the scale (and global) in the layout
the module holds — the container dtype, the swizzled artifact — since the loader runs the
quantization op after the converter ops.

Core loader: a parameter the quantizer's op is about to quantize is materialized in the checkpoint's
precision instead of the empty parameter's storage dtype (int8 storage zeroed FP4 weights; float8
double-rounded the existing FP8 path).

Also: `auto.py` groups the finegrained keys under one comment and drops unused NVFP4 imports; the
quantizer's config type hint names `FineGrainedConfig`; DeepGEMM guards (`_assert_affine_scales`,
`_assert_stacked_gate_up`, group-32 divisibility) run before the kernel load; expert biases take
the model dtype; comment trims for the noisy-comments check; tests for every path.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ion/Epilogue bundle fields

The kernels' dispatchers now take input_recipe/output_recipe strings, so the
bundle no longer needs the Quantization and Epilogue config classes; the
linears hand the module-level activation_format straight through as
input_recipe and the _kernel_quantization adapter is gone.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@vasqu

vasqu commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator

Wondering who's gonna review this 🌚

IlyasMoutawwakil and others added 3 commits September 13, 2026 03:48
…uantization/Epilogue bundle fields"

This reverts commit 07dae50. The kernels keep
the Quantization/Epilogue op-boundary configs as the dispatcher API, so the
bundle carries them and the linears build kernel.Quantization again.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…kip the kernels arch check

The kernels now speak the module's vocabulary: activation_format (None = the weights'
format, "bf16" = weight-only) goes straight to matmul_2d / matmul_grouped / the MoE
forwards, so the Quantization/Epilogue bundle fields and both adapter helpers are gone.

hub_kernels: kernels >= 0.16 checks a build's declared archs against the device; the
deep-gemm build declares only 9.0a although it JIT-compiles per device, so the mapping
opts out via a new check_arch passthrough (only forwarded when False, so older kernels
releases without the keyword keep working).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…(pr-1018)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

@vasqu vasqu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some initial comments from my side, sorry didnt have the power at the end of the core finegrained file anymore but I think I touched on the rest

Comment on lines +119 to +124
if tensor_idx is not None:
has_axis0_shard = any(self._normalize_param_dim(placement.dim) == 0 for _, placement in dim_placements)
if has_axis0_shard and not (
self._axis0_offset <= tensor_idx < self._axis0_offset + self._axis0_local_size
):
return None

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could be its own helper imo no wth the description. Imo we can shorten the comment, yes it happens on nvfp4 scales but imo this could become a global problem regardless no?

Comment thread src/transformers/integrations/deepgemm.py Outdated
Comment thread src/transformers/integrations/deepgemm.py Outdated
from .moe import ExpertsInterface, use_experts_implementation


warnings.warn(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

wouldnt this fire in all cases then? 🤔

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

rebump and potentially logger obj

"deep-gemm": {"repo_id": "kernels-community/deep-gemm", "version": 2},
# DeepGEMM JIT-compiles its CUDA kernels for the running device; the build metadata declares only
# the arch it was packaged on (9.0a), so the `kernels` arch check would wrongly reject sm_100.
"deep-gemm": {"repo_id": "kernels-community/deep-gemm", "version": 2, "check_arch": False},

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lol really? potentially different PR because that is worthy to fix regardless

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

rebump potentially

# DeepGEMM is CUDA-only, dynamic-only, SM90+, FP4 or 128x128-block FP8; the combos it would
# silently corrupt (float32 scales on SM100, #47030) are skipped up front rather than attempted.
# ``TRANSFORMERS_DISABLE_DEEPGEMM_LINEAR=1`` forces Triton for this dispatcher only.
deepgemm_preferred = (

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

wondering whether this bool should live as a fn in deepgeem instead

if self.has_bias:
self.bias = nn.Parameter(torch.empty(self.out_features))
else:
self.register_parameter("bias", None)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can be a helper function to set or register the parameter

Comment on lines +492 to +493
weight = to_local(self.weight)
scale_inv = to_local(self.weight_scale_inv)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is where the TP workaround finally comes around :D still think it should rather live on kernels side tbh or where we update our tp plans but it should be more "hidden"

Comment thread src/transformers/integrations/finegrained.py
Comment thread src/transformers/integrations/finegrained.py
IlyasMoutawwakil and others added 7 commits September 15, 2026 02:31
…s per-expert output norm

modelopt NVFP4 checkpoints ship a second-level global per projection and
a calibrated `input_scale`; both now reach the kernels, in either of the
two layouts such a checkpoint uses — one tensor per expert per projection
(GLM-5.2) or one stacked tensor per layer (the fused vLLM layout).

A gate|up stack arrives with two weight globals per expert, calibrated
separately. One conversion op owns every global of a layer and merges
them to the one-per-expert the kernels take: the stack keeps the gate's,
and the up half's folds onto the down projection, whose weight global
scales the expert output back up and whose calibrated input global moves
the other way, keeping the requantized intermediate on the range the
checkpoint calibrated. That is exact arithmetic on fp32 scalars, unlike
rescaling the up half's e4m3 block scales. Many-to-many converters are
now opt-in per op (`ConversionOps.supports_many_to_many`) rather than a
private whitelist.

A model states its per-expert output norm the way it states its
activation: `use_experts_implementation(post_expert_norm="<name>")` plus
an `_apply_post_norm` the class must define. A name the kernels implement
rides as that name and is folded into their reduce; anything else is the
module, called on the routed rows. The DeepGEMM experts paths, which have
no slot for such a norm, refuse instead of dropping it.

Weight-only modules hold no activation global, since nothing there
quantizes activations: the parameters are not allocated, the converters
for them are not emitted, and a calibrated checkpoint's `input_scale`
keys are ignored for that run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
One guard decorator on the DeepGEMM experts forwards instead of four statements
per body, and that file's comments cut to what is not already in the code.
`prefers_deepgemm_linear` moves to `deepgemm.py`, where its measured verdicts
belong, so the dispatcher just asks. `_FineGrainedModule.local(name)` is the one
accessor every forward reads operands through, so the DTensor unwrap appears
once rather than at seven call sites.

`_WeightFormat` is public (it is in `resolve_weight_format`'s signature), the
`TRANSFORMERS_FINEGRAINED_NO_SWIZZLE` debug hatch is gone (tests patch the
predicate), and `_set_optional_parameter` replaces six register-or-assign
blocks. `check_arch` is passed straight through: `kernels` is pinned >= 0.16,
where it landed, so the compatibility shim was dead.

The scalar-shard branch becomes `_owns_expert`, stated as the general fact it is
rather than as an NVFP4 anecdote. `FineGrainedConfig`'s docstring opens with a
format table. The frozen fp8 modules' `stacklevel=2` is dropped: at module scope
it named importlib rather than the module being deprecated.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…they served

Every checkpoint `quant_method` the two answer to — `fp8`, `mxfp8`, `mxfp4`,
`nvfp4`, `modelopt` — already routes to `FineGrainedHfQuantizer`, so both carry
the same DeprecationWarning the fp8 pair does.

Auditing that supersession found one real gap. `dequantize=True` on a GPT-OSS
MXFP4 checkpoint produced no converter at all: the `_blocks`/`_scales` keys were
wired only for the quantized path, so they went unmatched, where
`Mxfp4Config(dequantize=True)` had handled them. The finegrained chain now
dequantizes them, bit-equal to `mxfp4.convert_moe_packed_tensors` — the
implementation it replaces — which the new test compares against directly. Two
shared ops needed repair to get there: `FineGrainedPackedBlocks` regrouped a
sibling `_scales` entry as if it were blocks, and `FineGrainedDequantize` left
the scale it had consumed in the chain.

The same path on NVFP4 or modelopt would have folded one block scale and
dropped the second-level global, handing back a weight scaled by `1 / global`.
Neither frozen integration supported that either, but it failed quietly —
including where `dequantize` turns itself on (no GPU, compute capability < 8.9).
It raises now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
# Conflicts:
#	src/transformers/integrations/deepgemm.py
#	src/transformers/integrations/moe.py
It had become `FineGrainedFP8Config = FineGrainedConfig`, so the name still
imported but the class was gone. That makes "frozen for backward compatibility"
a fiction — the old name silently gained the new behaviour (`activation_format`,
the modelopt payload parsing, the wider `post_init`) — and it rendered the same
class under two `[[autodoc]]` headings. Restored verbatim, with a docstring line
pointing at the config that serves the whole family.

Splitting them means the two `isinstance` sites in `auto.py` no longer cover the
new config by accident. `LOADING_ATTRIBUTES_CONFIG_TYPES` is the load-bearing
one: it is what carries `dequantize` / `modules_to_not_convert` from a config
passed to `from_pretrained` onto the checkpoint's own, so leaving it alone would
have quietly stopped `FineGrainedConfig(dequantize=True)` from taking effect on
every fine-grained checkpoint. Both configs are in it now.

`quant_method="fp8"` still builds `FineGrainedConfig`: the frozen class is what
you get by naming it, not what a checkpoint resolves to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`c66564416e` restored the class and added a line saying it "configures the
frozen `finegrained_fp8` integration". That is not true: it sets
`quant_method="fp8"`, which `AUTO_QUANTIZER_MAPPING` routes to
`FineGrainedHfQuantizer` — the frozen quantizer is not reachable through
`from_pretrained` at all.

Removing it also leaves the class byte-identical to main, which is what a class
frozen for backward compatibility should look like in a diff. Which config
serves which family is already stated where a reader meets it, in
`FineGrainedConfig`'s own docstring on the same docs page.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@vasqu vasqu mentioned this pull request Sep 15, 2026
4 of 6 tasks
IlyasMoutawwakil and others added 4 commits September 16, 2026 01:10
Three CI failures on a CPU-only runner, three causes:

`post_expert_norm_name` is set by `use_experts_implementation`, but the shared
forwards are the ones every backend adapts onto — including duck-typed stand-ins
built without the decorator. The norm is optional, so read it as one.

`_assert_stacked_gate_up` already read its second attribute defensively; `has_gate`
now matches.

`prefers_deepgemm_linear` calls `is_sm100()`, and `get_device_capability()` asserts
rather than answering False on a CPU-only build. Caught there rather than gated on
`is_available()`, so that a test faking the capability alone still answers — which
is the convention the existing deepgemm tests are written to.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A converter whose target no module holds is a load failure, and the activation
global is exactly that under a weight-only run: the experts allocate one only when
they quantize activations, so the declarations are the only thing standing between
the two runs. Checked against a real module's parameter names in both.

Verified it catches the failure by removing the conditional: the weight-only run
then converts `gate_up_proj_input_global_scale` onto a module without that slot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The integration reads better as `is_available() and capability >= 10`; the reason it
wasn't was that three test sites faked the capability alone, which that form ignores.
That is the tests constraining the integration, so the tests move instead: they now
patch availability alongside the capability, and the module docstring says so.

Replaces the try/except that was catching what the availability check now prevents.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@IlyasMoutawwakil
IlyasMoutawwakil marked this pull request as ready for review September 16, 2026 09:45
IlyasMoutawwakil and others added 2 commits September 16, 2026 08:04
… name

A sharded parameter is already local by the time a forward runs, so the unwrap was
doing nothing: FSDP2 unshards in its pre-forward hook, and a TP/EP plan swaps each
DTensor for its local shard around the call (`MoeExpertsParallel` ->
`_use_local_dtensor_params`). Verified both: under `fully_shard` alone a param reads
as a plain Parameter inside forward, not a DTensor. So the invariant belongs to the
distributed layer that wrapped the parameter, not to every backend that reads one.

That removes `to_local` and the `local()` accessor. Absent slots still read back as
None because `_set_optional_parameter` registers them as None parameters, which is
what the kernels take for a missing bias or global.

`_moe_operands` also stops building its dict through a `both(suffix)` helper: the
kernels' arguments are now written out, so `down_proj_scale_inv=` can be found by
grep. The five `getattr`s that remain are load-bearing — the kernels always say
`gate_up_proj`, the module says `up_proj` when the model has no gate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The guard refused a per-expert output norm on every DeepGEMM arm, but only Mega MoE
has to: it fuses the routing-weighted reduce into its kernel, so there is no seam for
the norm. The other two arms reduce in `_combine_routed_output`, so the norm rides the
rows just before it — the same place the reference forwards apply it.

`deepgemm_experts_guards` gains `post_expert_norm`, which only the Mega MoE arm sets
False. A model naming a norm — Muse-Spark names `input_scaled_rms_norm` — was locked
out of every DeepGEMM path and now loses only Mega MoE.

The old behaviour had no test, which is why relaxing it broke nothing; the new one
covers both halves so the Mega MoE restriction cannot silently disappear either.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@vasqu vasqu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I went super into details this time, I feel like we should have good quality here as this is super impactful

Comment on lines 166 to 170

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Just noticed but we essentially have most of the checks here already no? we could introduce an elif for the not source ship case no?

@@ -578,6 +572,43 @@ def _combine_routed_output(


@deprecate_kwarg("output_dtype", version="v5.16")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can be removed now no?

def prefers_deepgemm_linear(
weight: torch.Tensor,
weight_scale_inv: torch.Tensor,
*,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
*,

hmm weird

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

rebump why do we need this

@@ -578,6 +572,43 @@ def _combine_routed_output(


@deprecate_kwarg("output_dtype", version="v5.16")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can remove this as well now

return decorate


@deepgemm_experts_guards()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit could probably refactored so we do not need the () if there is no kwarg added

Comment thread src/transformers/utils/quantization_config.py
Comment on lines +1752 to +1754
kwargs.pop("config_groups", None)
kwargs.pop("kv_cache_scheme", None)
kwargs.pop("producer", None)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

does this really hurt to keep?

prefix = glob[:-2] if glob.endswith(".*") else glob.rstrip("*").rstrip(".")
return re.escape(prefix) + r"(\..*)?$"

modules_to_not_convert = [_subtree_regex(g) for g in ignore]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

rebump, feel like this should become a general feature to allow regex and not be here from the get go

Comment thread src/transformers/utils/quantization_config.py Outdated
Comment thread src/transformers/core_model_loading.py Outdated
# Whether this op maps SEVERAL sources onto SEVERAL targets in one call — a converter is
# otherwise restricted to one-to-many, one-to-one or many-to-one, since an m:n mapping only
# makes sense when a single operation owns the whole relation (see `WeightConverter`).
supports_many_to_many: bool = False

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's just add onto _INTERNAL_MANY_TO_MANY_CONVERSIONS?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

would like to avoid introducing something new here

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the problem is that i have to add the finegrained conversion op to this constant so it has to be built dynamically

@ArthurZucker
ArthurZucker self-requested a review September 17, 2026 01:36
IlyasMoutawwakil and others added 6 commits September 17, 2026 06:47
A stacked experts projection is a Parameter, not a Module, so its scale
grid, bias and NVFP4 globals are SIBLINGS rather than children: the plan
matcher's one module-level fallback never reaches them, and the module
entry shards nothing. Each companion has to be named in the plan or the
weight becomes this rank's expert slice while its scales stay whole.

EP had a hand-written rule for exactly one of them (`grouped_gemm` ->
`_scale_inv`) that this branch had dropped; restore it and generalise to
`_bias`, both NVFP4 globals and a static activation scale. The gate_up
input global stays replicated: one per-tensor value for the pre-routing
hidden states.

TP never had the equivalent rule at all, since it uses packed_colwise /
rowwise and that loop only matched `grouped_gemm`. Add it, plus a second
fix: packed_colwise splits the output axis as two packed halves, which is
the [gate; up] STACK, while a quantized module holds the kernels'
INTERLEAVED rows. Register param-only styles for the axis actually wanted
-- the built-in styles count from the END of the shape, landing inside a
swizzled scale grid's inner tile and on a bias's expert axis.

Verified on 2 GPUs against a pre-quantized DeepSeek-V3 fixture: EP and TP
logits match an unsharded reference, and removing the TP half makes that
test fail with the scale left whole against a sharded weight. 8-rank smoke
runs pass on DeepSeek-V4-Flash (EP) and DeepSeek-V3 (TP); device_map smoke
passes on GLM-5.2-NVFP4, DeepSeek-V3, MiniMax-M3-MXFP8 and OLMoE.

Also in this pass, from the second review:
- `to_local` deprecated rather than deleted; every call site is gone now
  that `_hf_quantized_needs_local_tp` hands modules local tensors.
- `validate_environment` gains its first tests -- the only ones that
  existed built the frozen FP8 quantizer, which no `from_pretrained` can
  reach since AUTO_QUANTIZER_MAPPING["fp8"] became FineGrainedHfQuantizer.
- `_swizzles_scales` no longer claims block-FP8 never reaches a scaled
  MMA; it does, via UE8M0, but broadcasts one 128-block scalar in register
  so there is no per-group grid to lay out.
- modelopt remapped to nvfp4 at construction; glob skip-lists moved to
  the shared `subtree_pattern_to_regex`; GPT-OSS carries swiglu_alpha /
  swiglu_limit on its config.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…eights reach

Every `input_scale` rule lived in the modelopt converters, where the key means
NVFP4's second-level global. A plain-FP8 static checkpoint had no rule at all,
so Ministral-3 and Mistral-4 would have loaded with the registered default of
1.0 and quantized every activation against the wrong scale, silently.
`_static_activation_conversions` routes it to the slot each module holds: one
per expert for both MoE layouts, a rename for dense linears.
`FineGrainedInputGlobals` serves both meanings now and is named for what it
carries; only the NVFP4 gate_up global collapses to one value, because that
global is a split of the block scale and a static scale has no block level to
absorb an inflated one.

Adding the first per-tensor model to the load-path table then reached four
defects that nothing else does. A 0-dim scale was sharded as `Shard(ndim - 2)`
and had no axis-0 offset to take; the intra-expert rule split an `(E, 1, 1)`
scale on a length-1 axis; on-the-fly quantize emitted the dense NVFP4 global as
`weight_weight_global_scale` and left the expert input globals on meta, since
`ones_like` inherits the device of an unmaterialized slot; and an expert target
is a prefix of its own companions', so on save the weight converter claimed
`..._input_global_scale` and tried to split one value into gate|up halves.

The table and its round-trip assertion are newer than the last run of them:
`@slow` kept the suite that covers all of this out of every fast gate. Dropped
it — `@require_torch_multi_accelerator` already gates it on the hardware it
needs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Comment on lines +115 to +120
# A 0-dim tensor has no axis to shard, so every rank that owns it takes the whole value
# (`source[...]`, since a lazy safetensors slice rejects `source[()]`).
if not source_shape:
if tensor_idx is not None and not self._owns_expert(tensor_idx, dim_placements):
return None
return source[...].to(device=device, dtype=dtype)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Its weird we have to do this before, any reason we cannot do this later to enter the "moe" path later down below

if meta is None:
return
placement = Shard(1) if isinstance(module, torch.nn.Embedding) else Shard(meta.ndim - 2)
# a 0-dim scale is one value for the whole tensor: no axis to split, every rank keeps it

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

an example where this happens, ig some moe weights but interested to see

def prefers_deepgemm_linear(
weight: torch.Tensor,
weight_scale_inv: torch.Tensor,
*,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

rebump why do we need this

Comment on lines +691 to +707
def _assert_no_post_expert_norm(self: torch.nn.Module) -> None:
"""Mega MoE fuses the routing-weighted reduce into its kernel, so a model's per-expert output
norm has nowhere to go — it would be dropped silently, so refuse instead. The other DeepGEMM
arms reduce in `_combine_routed_output` and apply the norm on the rows just before it."""
if getattr(self, "has_post_expert_norm", False):
raise NotImplementedError(
"DeepGEMM Mega MoE cannot apply this model's per-expert output norm; use "
"`experts_implementation='deepgemm'`, 'grouped_mm' or 'batched_mm'."
)


def _apply_post_expert_norm(self: torch.nn.Module, rows: torch.Tensor) -> torch.Tensor:
"""The model's per-expert output norm, on the routed rows before the routing weights — the
same place the reference forwards apply it. A no-op for a model that declares none."""
if not getattr(self, "has_post_expert_norm", False):
return rows
return self._apply_post_norm(rows)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what happened here...?


def deepgemm_experts_guards(
forward=None,
*,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

same here, why do we have the * pattern here we shouldnt need it

Comment on lines +249 to +253
# Every companion beside a projection weight — block scales, bias, NVFP4 globals, a
# static activation scale — is expert-indexed too, so each shards with the weight. The
# matcher keys on the exact parameter name and falls back only to the owning MODULE,
# whose entry shards nothing, so without these the weight is this rank's expert slice
# while its scales are every rank's. gate_up's input global is per-tensor: replicated.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shorten

Comment on lines +276 to +297
# Intra-expert TP: the experts stay whole and the projection's own axis splits, so
# the per-expert globals, activation scales and the down bias (added after the
# row-reduce) stay replicated — but the scale grid still has to follow its weight.
# dim 1 is the output axis and dim 2 the input axis in BOTH the affine and the
# 5-D swizzled scale container, which is why naming those two suffices. A
# per-tensor scale has no such grid — one value per expert, `(E, 1, 1)`, whose
# axes are both 1 — so it joins the replicated companions and only its bias splits.
blocked = self.quantization_config.weight_block_size is not None
if projection == "down_proj" and style in ("rowwise", "packed_rowwise"):
if blocked:
updated_plan.setdefault(f"{key}_scale_inv", self._expert_shard(2))
elif projection != "down_proj" and style in ("packed_colwise", "colwise"):
# The companions take the WEIGHT's own split. The model declares which one it
# needs and is right: the interleave into the kernels' row order runs on each
# rank's shard, AFTER sharding, so the split has to match the layout the
# CHECKPOINT holds. `packed_colwise` is the strided split a concatenated
# `[gate; up]` keeps pairs together with; `colwise` the contiguous one a
# natively interleaved layout (GPT-OSS, `is_concatenated=False`) needs.
rows = self._expert_shard(1, packed=style == "packed_colwise")
if blocked:
updated_plan.setdefault(f"{key}_scale_inv", rows)
updated_plan.setdefault(f"{key}_bias", rows)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

shouldnt this be handled with the tp overrides instead? I would expect that tbh

per_expert = [
WeightConverter(
source_patterns=[
r"mlp\.experts\..*\.gate_proj\.weight_scale$",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why do we need the .*? we should expect unified layout

target_patterns=r"\1.activation_scale",
)
]
return per_expert + fused + dense

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any way we could just do something like in the prev fp8 where we did Fp8Dequantize only on all the scale patterns?

# "model.layers.1.*" and under `exclude_modules` bare ("self_attn", "layers.0.")
from ..quantizers.quantizers_utils import subtree_pattern_to_regex

modules_to_not_convert = [subtree_pattern_to_regex(g) for g in ignore]

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ah ic but shouldnt we use ours instead, it still feels like its separate PR as it could be a general feat to allow regex

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

then we do not need this but modules to convert can just take it as is

IlyasMoutawwakil and others added 7 commits September 21, 2026 05:14
Reload: an on-the-fly quantized model lost its swizzle reverse, so saving one
wrote swizzled scales a reload could not read back. keep_swizzle_reverse_for_save
retains it, and the post-load pass is split into the two delegations it always
was, with `covered` computed from source-pattern matches against a fused probe
name rather than assumed.

TP: a tied lm_head IS the input embedding's parameter, so the embedding_rowwise
entry that PretrainedConfig adds for tied weights shards the tensor lm_head
reads, while lm_head itself sits outside base_model_tp_plan and gets no style —
leaving F.linear with a plain input against a DTensor weight. FSDP resolves the
pair; TP had nothing. init_parallel_plans now gives it colwise_rep. It reads the
config flag rather than `head.weight is embed.weight` because tying happens
later in post_init, where the identity test silently misses.

glm4v_moe shipped no expert entries, so its experts were never sharded. Added
packed_colwise / rowwise / moe_tp_experts plus shared_experts.

DeepseekV4RMSNorm returns the input dtype. V4 pins these norms to FP32 via
_keep_in_fp32_modules_strict, so `weight * hidden_states` promoted the result
and handed FP32 to compressor projections that ship BF16. The cast-only form is
bit-identical where the weight already matches (torch.equal).

Tests: pin the fixture to bf16, which exposed three defects (two of them
upstream, reproducing with no quantizer at all); add a rounding floor that
perturbs every sharded block by one ULP, so the load-path comparison measures
against accumulated reduction-order drift instead of an absolute tolerance;
size shards against the checkpoint rather than against the other leg.
_checkpoint_expert_shape also learns DeepSeek's w1/w3 spelling — it only knew
gate_proj/up_proj and the fused layout, so the V4 fp4-experts fixture matched
neither and failed outright.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`update_tp_plan` carried `if "Qwen3" in config.__class__.__name__:` and then
ASSIGNED a hand-written dense plan over `config.base_model_tp_plan`. Three things
were wrong with it:

- redundant: the module-keyed entries Qwen3 already ships reach a parameter's
  scale through the plan matcher's module fallback — `q_proj.weight_scale_inv`
  resolves to `colwise`, `down_proj.weight_scale_inv` to `rowwise` — so the
  parameter-keyed duplicates bought nothing;
- destructive: assigning dropped what the model ships, including `q_norm`/`k_norm`
  `replicated_with_grad_allreduce`, and on Qwen3-MoE the whole expert plan
  (`experts: moe_tp_experts`, `experts.gate_up_proj: packed_colwise`, ...). The
  companion pass immediately below keys on `moe_tp_experts`, so it then found
  nothing and the expert scales got no entry either;
- over-broad: 25 config classes contain "Qwen3". On the multimodal ones, whose
  plans live on a sub-config and whose outer config ships none, it stamped a dense
  text plan onto the outer config where those paths match nothing.

Removed. Qwen3 and Qwen3-MoE now keep every shipped entry, Qwen3-MoE additionally
gains the three expert companions it was being denied, and a multimodal config's
outer plan stays empty. Three assertions cover it, against REAL configs — the
existing plan tests build a SimpleNamespace, which a branch keyed on the config
CLASS is structurally invisible to.

Structure:
- the conversion ops move to `integrations/finegrained_conversions.py`
  (finegrained.py 1631 -> 985). The dependency is one-way;
  `keep_swizzle_reverse_for_save` goes with them, which is what removes the cycle
  rather than needing a local import to break it.
- `_nvfp4_conversions` splits per checkpoint layout. The two never interact — a
  checkpoint matches one set and the other never fires — so they were adjacent,
  not related.
- `replace_with_finegrained_layer` hoists the kwargs every swap shares, leaving
  each branch showing only what is its own.

Review comments:
- TP styles are registered properly (`moe_experts_colwise` /
  `moe_experts_packed_colwise` / `moe_experts_rowwise`) instead of
  `ALL_PARALLEL_STYLES.register` at plan-build time.
- the frozen-fp8 deprecation notices fired at IMPORT, for anyone importing
  `transformers.integrations` at all, one of them through `warnings.warn` rather
  than the logger. Now one `logger.warning_once` on quantizer construction.
- `has_global_scale` dropped: `global_scale_dtype: torch.dtype | None` already
  answered it.
- defensive `getattr` on declared config fields replaced with direct access.
- `quantizers_utils.py`, `integrations/tensor_parallel.py` and
  `integrations/finegrained_fp8.py` are back to their upstream contents and leave
  this diff entirely.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The ten conversion ops were five self-inverting (reverse is themselves with the
direction flipped) and five paired (reverse names a different class), and every
one of them spelled out its own `__init__` and `reverse_op`. A single
`_FineGrainedOp` base carries the quantizer and defaults `reverse_op` to
`type(self)(..., inverse=not self.inverse)`; the paired ops override it and
ignore the flag. Twelve lines replace fifty-two, and `type(self)` cannot go stale
the way five hardcoded class names could.

- `_moe_operands`: the twelve projection entries were `{kernel_name:
  getattr(module, held_name)}` twelve times over. One comprehension states the
  rule instead — the kernels always say `gate_up_proj`, the module says
  `up_proj` where the model has no gate.
- `FineGrainedExperts.__init__` (104 lines) was reading config and allocating
  parameters in one breath. Allocation moves to `_register_projection`, called
  once per projection, with the four `_set_optional_parameter` calls folded into
  a loop and two loop-invariants hoisted out.
- `_with_expert_layout_ops`: a six-line comment explaining a condition becomes
  `sources_the_fused_key`, a named predicate carrying that reasoning; the loop
  body goes from fourteen lines to three. The expert-target regex compiles once.
- the `expert_dtype` comment now says why the side-channel exists: DeepSeek-V4 is
  MIXED (mxfp4 experts under an "fp8" quant config) and a quant config carries a
  single `quant_method` with no way to express it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two branches decided behaviour by matching a NAME, which is how the Qwen3 plan
override stayed invisible — the default arm silently absorbs a miss:

- the grouped-linear swap tested `"GroupedLinear" in type(module).__name__`. It
  now tests `hasattr(module, "n_groups")`, the attribute the branch consumes, so
  a grouped linear named anything else is still swapped and one named right but
  shaped wrong no longer reaches an AttributeError. ConvBERT's
  `GroupedLinearLayer` was never caught (it is an `nn.Module`), so this is
  hardening rather than a fix.
- `_global_role` classified with `"w2" in key`, an unanchored substring for what
  is a path SEGMENT. Since the default is gate_up, a stray match merged a down
  global into the wrong bucket silently. Matched on segments now.

Readability:

- `FineGrainedQuantize` and `FineGrainedDequantize` still carried their own
  `__init__`; the base-class collapse only matched the form with a default.
- `_weight_holder`'s pair was indexed as `holder[0]` / `holder[1]` in four
  places, and is now unpacked into `module, scale_name`.
- `_recontain` -> `_as_container` and `_swizzles_scales` ->
  `_holds_swizzled_scales`: both were coined, and the codebase already says
  "container" (`scale_container_dtype`) and "holds" (`holds_interleaved_gate_up`).
- `q` / `s` -> `blocked` / `per_block` in the dequant reshape.
- each file opens with a short statement of what it is and where its neighbour
  lives, rather than none at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Groups: every quant config is GROUPED internally, and the flat fields a
single-format checkpoint ships normalise into one catch-all group. A mixed
checkpoint needs more than one -- DeepSeek-V4 is W4A4 mxfp4 EXPERTS over
block-FP8 linears -- and compressed-tensors/modelopt already describe that
parametrically under `config_groups`, which was being dropped. `group_for(name)`
answers with the group that owns a module, so the swap stops asking the config
which format the whole model is. A targeted group wins over the catch-all
whatever order they were declared in, because `to_json_string` sorts keys and a
config that round-tripped through `config.json` has lost that order.
`split_out_experts` folds the legacy `config.expert_dtype` side-channel into the
two groups it always meant.

Four defects, found bisecting a red load-path suite:

- the experts' static `activation_scale` was allocated without an explicit
  dtype, so `from_pretrained` gave it the checkpoint's bf16. Its two siblings
  kept theirs; a loop collapse made the odd one out invisible.
- that scale was never WRITTEN. The loader materialises a key no checkpoint
  supplies with `torch.empty_like` and `_init_weights` has no branch for a
  scale, so it reached the kernels as uninitialised memory -- read back as 0 and
  -0.0244 on a single device. A zero divides the activations by zero, which is
  where the NaN logits came from, and the randomness is why they moved between
  runs. `FineGrainedQuantize` now writes every calibration-fed slot as the
  identity, which is what it already did for the NVFP4 global.
- `_quantize_block_fp8` refused a shape the block does not divide. The format's
  own producers do not: DeepSeek-V3 ships `kv_a_proj_with_mqa` as (576, 7168)
  E4M3 against a 128x128 block with a (5, 56) scale grid. It pads the trailing
  partial block now and rounds the grid up to match, so on-the-fly quantization
  lands what the checkpoint would have.
- that refusal was doubling as the filter for tensors that are not weights: an
  expert bias is 2-D and `E % 128` never divided, so it fell out there. Turned
  away against `_weight_holder` instead, which already knew.

With shapes no longer declined, a full-precision weight in a quantized slot can
only mean the checkpoint does not quantize a module it targets -- and the
forward has no path for it -- so the three `element_size() > 1` fallbacks go and
the check raises rather than warns.

Tests: dtype invariance under both default dtypes (rather than asserting fp32,
so it catches the next slot too), identity activation scales, ceil-rounded scale
grids, and bias pass-through. Each verified to fail without its fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- `_quantize_block_fp8` loses both `if pad_m or pad_n` branches: padding by
  `-rows % block_m` is already a no-op when the shape tiles, and cropping back
  is then a full slice, which `contiguous` returns unchanged.
- `_identity_activation_scales` names the projection prefix once instead of
  probing `hasattr` per slot -- a dense linear's stem IS `weight`, which is
  what distinguishes it from the experts' `<proj>_`.
- the two new comment blocks say the same thing in half the lines.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`FineGrainedGroupedLinear` took `block_size`/`activation_scheme`/`scale_fmt` but
not `weight_format`/`activation_format`, so it silently built its parent's fp8
default -- right for DeepSeek-V4's block-FP8 grouped linears today, wrong the
moment a group says otherwise. It takes the pair now, which is also what let the
swap drop the comprehension that picked three keys back out of `storage_for`:
all three branches pass the same mapping.

The experts branch was 27 of the dispatch's 40 lines, so the dispatch itself was
the hard part to find. `_quantized_experts` holds the flags-twice-over dance and
the post-norm carry-over, and the three-way choice now reads at a glance.

`FineGrainedQuantize`'s docstring still said shapes that don't tile pass
through, which stopped being true when the block quantizer learned to pad. It
names `_weight_holder` as what actually decides, since rank does not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

CI recap

Dashboard: View test results in Grafana
Latest run: 35342495884:1
Result: failure | Jobs: 16 | Tests: 190,675 | Failures: 1 | Duration: 13h 28m

Code quality check failed: test jobs were skipped. Fix the code quality issues and push again to run tests.

@github-actions

Copy link
Copy Markdown
Contributor

[For maintainers] Suggested jobs to run (before merge)

run-slow: deepseek_v4, glm4v_moe, gpt_oss, openai_privacy_filter, finegrained_fp8, mxfp4, nvfp4

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants