Repository navigation
Finegrained kernels: unified quantization kernels api - #48058
IlyasMoutawwakil wants to merge 44 commits into
Conversation
The kernels now store gate|up row-interleaved (gate j at row 2j, up j at 2j+1)
rather than as two N-apart halves, so the integration follows:
- GPT-OSS already ships this order on disk, so its de-interleave converters go
away entirely: _deinterleave_gate_up_rows, its use in FineGrainedMxfp4Deserialize,
and the whole FineGrainedGateUpBiasDeinterleave op plus its registration.
- Checkpoints that ship gate_up stacked are interleaved once post-load by
interleave_gate_up_after_loading, ordered BEFORE swizzle_scales_after_loading
(the swizzle packs whatever row order it finds) and skipped for mxfp4. This
cannot live in conversion_mapping's MergeModulelist+Concatenate chain — that is
shared with non-finegrained MoE models.
- _apply_gate de-interleaves with a stride-2 split, mirroring the fused epilogue's
split_gate_up; the two must agree or fused and unfused stop being comparable.
- swizzle_scales_after_loading drops gate= and gates on the doubled extent
(n_rows % 128, i.e. N % 64), so GPT-OSS N=2880 now pre-swizzles instead of
falling back to an affine scale read.
Adds tests/kernels/test_finegrained.py (24 tests), covering the post-load interleave,
that mxfp4 is left untouched, dispatch/swizzle gating, and the MergeModulelist path.
tests/kernels/test_finegrained.py: 24 passed.
The kernels store gate|up row-interleaved, but ``transform_weights_for_mega_moe`` does its OWN gate/up interleave and so takes the stacked form. Interleaving those modules at load and undoing it at the boundary is a pointless round trip — and interleaving twice is not the identity, it is a different permutation, i.e. silent garbage rather than a crash. So modules bound for megamoe are simply never interleaved: the post-load pass skips them and records what the module actually holds in ``_gate_up_interleaved`` (the bytes cannot be inspected to tell, and ``set_experts_implementation`` can switch backends after load). ``setup_megamoe_weights`` reads that flag and raises if it is ever handed an interleaved module, rather than re-permuting on an assumption. The other DeepGEMM expert paths need no change: their grouped GEMM is layout-agnostic and the gate split happens in ``_apply_gate``, which already de-interleaves stride-2. Verified: triton backend interleaves and flags; megamoe backend is left untouched with no round trip. The raise itself is unverified here — DeepGEMM's JIT needs a CUDA toolkit >= 12.9 and CUDA_HOME is unset on this box, so no megamoe path executes.
|
The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update. |
# Conflicts: # src/transformers/integrations/hub_kernels.py # src/transformers/quantizers/auto.py
The memoized `(epilogue, gate_up_quantization, down_quantization)` triple bought nothing where it matters. Constructing the three dataclasses measured ~1.4us per layer (~85us per step at 61 layers), and only in EAGER decode — under cudagraphs the host path does not run on replay, which is the mode this integration deploys in. Against that it held hidden mutable state on the module, and cached a gate_up quantization carrying an `output_recipe` that silently mismatches any caller asking for the unfused form. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The kernels now take a per-expert bias as a `bias=` operand, so a biased model no longer falls back to the unfused two-GEMM form just for having a bias. `_kernel_epilogue` stops bailing on `has_bias`, and both host-side adds go away: `_apply_unfused_gate_up` keeps only the activation, `_finish_down` only the routing-weighted reduce. That also drops the `torch.sort` the grouped path paid to index the gate_up bias by expert-sorted row -- the kernel indexes by the tile's own expert id instead. An unsupported act_fn no longer costs the bias its fusion either: the bias is an operand beside the epilogue rather than part of it, so it fuses whether or not the GLU does, and `epilogue is None` goes back to meaning exactly one thing. The bias rides the same output-row axis as the weight and scale grid, so `interleave_gate_up_after_loading` already delivers it in the kernels' interleaved order -- no reordering here. Tests assert the operand reaches both GEMMs, plus a new case pinning that an unfusable activation still fuses its bias. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Conflict resolutions: - quantizers/auto.py: keep the finegrained claims on fp8/nvfp4/modelopt (pre-quantized + MoE experts; NVFP4HfQuantizer only quantizes on the fly), take main's new gguf entry. - integrations/__init__.py: union — the finegrained module plus main's FP8Embedding exports. - integrations/deepgemm.py: ours (to_local unwrap + the stacked-gate guard); re-add the to_local import main's shim conversion dropped. - integrations/tensor_parallel.py: main's shim, with to_local restored on it — the kernel integrations still import it from this path. - distributed/sharding_utils.py: re-apply the 0-dim scalar guard (ModelOpt per-projection weight_scale_2/input_scale) inside the rewritten DtensorShardOperation.shard_tensor: replicated on the dense path, expert-ownership-filtered on the MoE path. - tests/tensor_parallel: main's rewrite, with the 0-dim sharding test re-added against the new engine (passes). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…e kernels' MoE chain
The experts forwards are adapters over the kernels' `moe_fused_batched/grouped`: `_moe_operands`
hands them the module's tensors and the activation — a `get_supported_act_fns()` name the gate_up
epilogue fuses, else the module's own GLU as a callable the kernels run on the host, so any
activation works without a kernel release. The kernel bundle is loaded by symbol name
(`matmul_2d/batched/grouped`, the two forwards, swizzle/unswizzle, `get_supported_act_fns`).
Every layout difference between a checkpoint and what the modules hold is a `ConversionOps` with a
reverse, attached by the quantizer, so `save_pretrained` restores the checkpoint bitwise:
- `FineGrainedInterleaveGateUp`: stacked [gate; up] rows -> the kernels' [g0, u0, ...] order,
skipped when the model declares `is_concatenated=False` (GPT-OSS; the flag travels through
`use_experts_implementation`, which stamps its defaults after `__init__`) or the backend packs
gate|up itself (DeepGEMM Mega MoE);
- `FineGrainedScaleContainer`: a scale into the dtype its module holds — the same bytes for a
uint8 UE8M0 container (MiniMax), an exact cast for float32 values (dsv4-flash-base) — and the
module records the shipped container so the reverse restores it;
- `FineGrainedSwizzleScales`: the SWIZZLE_32_4_4 artifact for modules that hold 5-D scales
(`FineGrainedExperts.__init__` allocates that shape for every non-DeepGEMM backend on SM100
with a group-scaled format and activation quant; whole 128-row/4-col blocks only);
- `FineGrainedPackedBlocks` regroups GPT-OSS `{proj}_blocks` into the packed weight, so the
blocks/scales format is two one-to-one converters and round-trips too.
Catch-all converters cover keys that already arrive under the fused names and dense scale keys.
The post-load hooks, `_save_to_state_dict` override and post-load dtype cast are gone.
Also: the eager per-expert loop reads a swizzled stack's `[e:e+1]` slice and passes bias, NVFP4
global and activation format (it dropped all three); expert biases take the model dtype; the
`mxfp4` format pins UE8M0 scales; `FineGrainedEmbedding` (FP8 table, per-tensor scale) for
`modules_to_convert`; DeepGEMM backends refuse modules holding swizzled scales or interleaved rows
before loading the kernel; `disable_deepgemm_on_multi_device` is public.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…heckpoint precision for quantized params `FineGrainedQuantize` reads the module's weight format: block-FP8 in torch (per-block E4M3, fp32 or power-of-two UE8M0 inverse scales, from the module's block and held scale dtype), MXFP8 / MXFP4 / NVFP4 through the kernels' row-wise quantizers in one launch per tensor (NVFP4 normalized by the per-tensor / per-expert global `amax / (6 * 448)`), and emits the scale (and global) in the layout the module holds — the container dtype, the swizzled artifact — since the loader runs the quantization op after the converter ops. Core loader: a parameter the quantizer's op is about to quantize is materialized in the checkpoint's precision instead of the empty parameter's storage dtype (int8 storage zeroed FP4 weights; float8 double-rounded the existing FP8 path). Also: `auto.py` groups the finegrained keys under one comment and drops unused NVFP4 imports; the quantizer's config type hint names `FineGrainedConfig`; DeepGEMM guards (`_assert_affine_scales`, `_assert_stacked_gate_up`, group-32 divisibility) run before the kernel load; expert biases take the model dtype; comment trims for the noisy-comments check; tests for every path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ion/Epilogue bundle fields The kernels' dispatchers now take input_recipe/output_recipe strings, so the bundle no longer needs the Quantization and Epilogue config classes; the linears hand the module-level activation_format straight through as input_recipe and the _kernel_quantization adapter is gone. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Wondering who's gonna review this 🌚 |
…uantization/Epilogue bundle fields" This reverts commit 07dae50. The kernels keep the Quantization/Epilogue op-boundary configs as the dispatcher API, so the bundle carries them and the linears build kernel.Quantization again. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…kip the kernels arch check The kernels now speak the module's vocabulary: activation_format (None = the weights' format, "bf16" = weight-only) goes straight to matmul_2d / matmul_grouped / the MoE forwards, so the Quantization/Epilogue bundle fields and both adapter helpers are gone. hub_kernels: kernels >= 0.16 checks a build's declared archs against the device; the deep-gemm build declares only 9.0a although it JIT-compiles per device, so the mapping opts out via a new check_arch passthrough (only forwarded when False, so older kernels releases without the keyword keep working). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
3bab7d9 to
240abc9
Compare
…(pr-1018) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
vasqu
left a comment
There was a problem hiding this comment.
Some initial comments from my side, sorry didnt have the power at the end of the core finegrained file anymore but I think I touched on the rest
| if tensor_idx is not None: | ||
| has_axis0_shard = any(self._normalize_param_dim(placement.dim) == 0 for _, placement in dim_placements) | ||
| if has_axis0_shard and not ( | ||
| self._axis0_offset <= tensor_idx < self._axis0_offset + self._axis0_local_size | ||
| ): | ||
| return None |
There was a problem hiding this comment.
Could be its own helper imo no wth the description. Imo we can shorten the comment, yes it happens on nvfp4 scales but imo this could become a global problem regardless no?
| from .moe import ExpertsInterface, use_experts_implementation | ||
|
|
||
|
|
||
| warnings.warn( |
There was a problem hiding this comment.
wouldnt this fire in all cases then? 🤔
There was a problem hiding this comment.
rebump and potentially logger obj
| "deep-gemm": {"repo_id": "kernels-community/deep-gemm", "version": 2}, | ||
| # DeepGEMM JIT-compiles its CUDA kernels for the running device; the build metadata declares only | ||
| # the arch it was packaged on (9.0a), so the `kernels` arch check would wrongly reject sm_100. | ||
| "deep-gemm": {"repo_id": "kernels-community/deep-gemm", "version": 2, "check_arch": False}, |
There was a problem hiding this comment.
lol really? potentially different PR because that is worthy to fix regardless
| # DeepGEMM is CUDA-only, dynamic-only, SM90+, FP4 or 128x128-block FP8; the combos it would | ||
| # silently corrupt (float32 scales on SM100, #47030) are skipped up front rather than attempted. | ||
| # ``TRANSFORMERS_DISABLE_DEEPGEMM_LINEAR=1`` forces Triton for this dispatcher only. | ||
| deepgemm_preferred = ( |
There was a problem hiding this comment.
wondering whether this bool should live as a fn in deepgeem instead
| if self.has_bias: | ||
| self.bias = nn.Parameter(torch.empty(self.out_features)) | ||
| else: | ||
| self.register_parameter("bias", None) |
There was a problem hiding this comment.
can be a helper function to set or register the parameter
| weight = to_local(self.weight) | ||
| scale_inv = to_local(self.weight_scale_inv) |
There was a problem hiding this comment.
this is where the TP workaround finally comes around :D still think it should rather live on kernels side tbh or where we update our tp plans but it should be more "hidden"
…s per-expert output norm modelopt NVFP4 checkpoints ship a second-level global per projection and a calibrated `input_scale`; both now reach the kernels, in either of the two layouts such a checkpoint uses — one tensor per expert per projection (GLM-5.2) or one stacked tensor per layer (the fused vLLM layout). A gate|up stack arrives with two weight globals per expert, calibrated separately. One conversion op owns every global of a layer and merges them to the one-per-expert the kernels take: the stack keeps the gate's, and the up half's folds onto the down projection, whose weight global scales the expert output back up and whose calibrated input global moves the other way, keeping the requantized intermediate on the range the checkpoint calibrated. That is exact arithmetic on fp32 scalars, unlike rescaling the up half's e4m3 block scales. Many-to-many converters are now opt-in per op (`ConversionOps.supports_many_to_many`) rather than a private whitelist. A model states its per-expert output norm the way it states its activation: `use_experts_implementation(post_expert_norm="<name>")` plus an `_apply_post_norm` the class must define. A name the kernels implement rides as that name and is folded into their reduce; anything else is the module, called on the routed rows. The DeepGEMM experts paths, which have no slot for such a norm, refuse instead of dropping it. Weight-only modules hold no activation global, since nothing there quantizes activations: the parameters are not allocated, the converters for them are not emitted, and a calibrated checkpoint's `input_scale` keys are ignored for that run. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
One guard decorator on the DeepGEMM experts forwards instead of four statements per body, and that file's comments cut to what is not already in the code. `prefers_deepgemm_linear` moves to `deepgemm.py`, where its measured verdicts belong, so the dispatcher just asks. `_FineGrainedModule.local(name)` is the one accessor every forward reads operands through, so the DTensor unwrap appears once rather than at seven call sites. `_WeightFormat` is public (it is in `resolve_weight_format`'s signature), the `TRANSFORMERS_FINEGRAINED_NO_SWIZZLE` debug hatch is gone (tests patch the predicate), and `_set_optional_parameter` replaces six register-or-assign blocks. `check_arch` is passed straight through: `kernels` is pinned >= 0.16, where it landed, so the compatibility shim was dead. The scalar-shard branch becomes `_owns_expert`, stated as the general fact it is rather than as an NVFP4 anecdote. `FineGrainedConfig`'s docstring opens with a format table. The frozen fp8 modules' `stacklevel=2` is dropped: at module scope it named importlib rather than the module being deprecated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…they served Every checkpoint `quant_method` the two answer to — `fp8`, `mxfp8`, `mxfp4`, `nvfp4`, `modelopt` — already routes to `FineGrainedHfQuantizer`, so both carry the same DeprecationWarning the fp8 pair does. Auditing that supersession found one real gap. `dequantize=True` on a GPT-OSS MXFP4 checkpoint produced no converter at all: the `_blocks`/`_scales` keys were wired only for the quantized path, so they went unmatched, where `Mxfp4Config(dequantize=True)` had handled them. The finegrained chain now dequantizes them, bit-equal to `mxfp4.convert_moe_packed_tensors` — the implementation it replaces — which the new test compares against directly. Two shared ops needed repair to get there: `FineGrainedPackedBlocks` regrouped a sibling `_scales` entry as if it were blocks, and `FineGrainedDequantize` left the scale it had consumed in the chain. The same path on NVFP4 or modelopt would have folded one block scale and dropped the second-level global, handing back a weight scaled by `1 / global`. Neither frozen integration supported that either, but it failed quietly — including where `dequantize` turns itself on (no GPU, compute capability < 8.9). It raises now. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
# Conflicts: # src/transformers/integrations/deepgemm.py # src/transformers/integrations/moe.py
It had become `FineGrainedFP8Config = FineGrainedConfig`, so the name still imported but the class was gone. That makes "frozen for backward compatibility" a fiction — the old name silently gained the new behaviour (`activation_format`, the modelopt payload parsing, the wider `post_init`) — and it rendered the same class under two `[[autodoc]]` headings. Restored verbatim, with a docstring line pointing at the config that serves the whole family. Splitting them means the two `isinstance` sites in `auto.py` no longer cover the new config by accident. `LOADING_ATTRIBUTES_CONFIG_TYPES` is the load-bearing one: it is what carries `dequantize` / `modules_to_not_convert` from a config passed to `from_pretrained` onto the checkpoint's own, so leaving it alone would have quietly stopped `FineGrainedConfig(dequantize=True)` from taking effect on every fine-grained checkpoint. Both configs are in it now. `quant_method="fp8"` still builds `FineGrainedConfig`: the frozen class is what you get by naming it, not what a checkpoint resolves to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`c66564416e` restored the class and added a line saying it "configures the frozen `finegrained_fp8` integration". That is not true: it sets `quant_method="fp8"`, which `AUTO_QUANTIZER_MAPPING` routes to `FineGrainedHfQuantizer` — the frozen quantizer is not reachable through `from_pretrained` at all. Removing it also leaves the class byte-identical to main, which is what a class frozen for backward compatibility should look like in a diff. Which config serves which family is already stated where a reader meets it, in `FineGrainedConfig`'s own docstring on the same docs page. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Three CI failures on a CPU-only runner, three causes: `post_expert_norm_name` is set by `use_experts_implementation`, but the shared forwards are the ones every backend adapts onto — including duck-typed stand-ins built without the decorator. The norm is optional, so read it as one. `_assert_stacked_gate_up` already read its second attribute defensively; `has_gate` now matches. `prefers_deepgemm_linear` calls `is_sm100()`, and `get_device_capability()` asserts rather than answering False on a CPU-only build. Caught there rather than gated on `is_available()`, so that a test faking the capability alone still answers — which is the convention the existing deepgemm tests are written to. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A converter whose target no module holds is a load failure, and the activation global is exactly that under a weight-only run: the experts allocate one only when they quantize activations, so the declarations are the only thing standing between the two runs. Checked against a real module's parameter names in both. Verified it catches the failure by removing the conditional: the weight-only run then converts `gate_up_proj_input_global_scale` onto a module without that slot. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The integration reads better as `is_available() and capability >= 10`; the reason it wasn't was that three test sites faked the capability alone, which that form ignores. That is the tests constraining the integration, so the tests move instead: they now patch availability alongside the capability, and the module docstring says so. Replaces the try/except that was catching what the availability check now prevents. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… name A sharded parameter is already local by the time a forward runs, so the unwrap was doing nothing: FSDP2 unshards in its pre-forward hook, and a TP/EP plan swaps each DTensor for its local shard around the call (`MoeExpertsParallel` -> `_use_local_dtensor_params`). Verified both: under `fully_shard` alone a param reads as a plain Parameter inside forward, not a DTensor. So the invariant belongs to the distributed layer that wrapped the parameter, not to every backend that reads one. That removes `to_local` and the `local()` accessor. Absent slots still read back as None because `_set_optional_parameter` registers them as None parameters, which is what the kernels take for a missing bias or global. `_moe_operands` also stops building its dict through a `both(suffix)` helper: the kernels' arguments are now written out, so `down_proj_scale_inv=` can be found by grep. The five `getattr`s that remain are load-bearing — the kernels always say `gate_up_proj`, the module says `up_proj` when the model has no gate. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The guard refused a per-expert output norm on every DeepGEMM arm, but only Mega MoE has to: it fuses the routing-weighted reduce into its kernel, so there is no seam for the norm. The other two arms reduce in `_combine_routed_output`, so the norm rides the rows just before it — the same place the reference forwards apply it. `deepgemm_experts_guards` gains `post_expert_norm`, which only the Mega MoE arm sets False. A model naming a norm — Muse-Spark names `input_scaled_rms_norm` — was locked out of every DeepGEMM path and now loses only Mega MoE. The old behaviour had no test, which is why relaxing it broke nothing; the new one covers both halves so the Mega MoE restriction cannot silently disappear either. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
vasqu
left a comment
There was a problem hiding this comment.
I went super into details this time, I feel like we should have good quality here as this is super impactful
There was a problem hiding this comment.
Just noticed but we essentially have most of the checks here already no? we could introduce an elif for the not source ship case no?
| @@ -578,6 +572,43 @@ def _combine_routed_output( | |||
|
|
|||
|
|
|||
| @deprecate_kwarg("output_dtype", version="v5.16") | |||
| def prefers_deepgemm_linear( | ||
| weight: torch.Tensor, | ||
| weight_scale_inv: torch.Tensor, | ||
| *, |
There was a problem hiding this comment.
| *, |
hmm weird
| @@ -578,6 +572,43 @@ def _combine_routed_output( | |||
|
|
|||
|
|
|||
| @deprecate_kwarg("output_dtype", version="v5.16") | |||
There was a problem hiding this comment.
can remove this as well now
| return decorate | ||
|
|
||
|
|
||
| @deepgemm_experts_guards() |
There was a problem hiding this comment.
nit could probably refactored so we do not need the () if there is no kwarg added
| kwargs.pop("config_groups", None) | ||
| kwargs.pop("kv_cache_scheme", None) | ||
| kwargs.pop("producer", None) |
There was a problem hiding this comment.
does this really hurt to keep?
| prefix = glob[:-2] if glob.endswith(".*") else glob.rstrip("*").rstrip(".") | ||
| return re.escape(prefix) + r"(\..*)?$" | ||
|
|
||
| modules_to_not_convert = [_subtree_regex(g) for g in ignore] |
There was a problem hiding this comment.
rebump, feel like this should become a general feature to allow regex and not be here from the get go
| # Whether this op maps SEVERAL sources onto SEVERAL targets in one call — a converter is | ||
| # otherwise restricted to one-to-many, one-to-one or many-to-one, since an m:n mapping only | ||
| # makes sense when a single operation owns the whole relation (see `WeightConverter`). | ||
| supports_many_to_many: bool = False |
There was a problem hiding this comment.
Let's just add onto _INTERNAL_MANY_TO_MANY_CONVERSIONS?
There was a problem hiding this comment.
would like to avoid introducing something new here
There was a problem hiding this comment.
the problem is that i have to add the finegrained conversion op to this constant so it has to be built dynamically
A stacked experts projection is a Parameter, not a Module, so its scale grid, bias and NVFP4 globals are SIBLINGS rather than children: the plan matcher's one module-level fallback never reaches them, and the module entry shards nothing. Each companion has to be named in the plan or the weight becomes this rank's expert slice while its scales stay whole. EP had a hand-written rule for exactly one of them (`grouped_gemm` -> `_scale_inv`) that this branch had dropped; restore it and generalise to `_bias`, both NVFP4 globals and a static activation scale. The gate_up input global stays replicated: one per-tensor value for the pre-routing hidden states. TP never had the equivalent rule at all, since it uses packed_colwise / rowwise and that loop only matched `grouped_gemm`. Add it, plus a second fix: packed_colwise splits the output axis as two packed halves, which is the [gate; up] STACK, while a quantized module holds the kernels' INTERLEAVED rows. Register param-only styles for the axis actually wanted -- the built-in styles count from the END of the shape, landing inside a swizzled scale grid's inner tile and on a bias's expert axis. Verified on 2 GPUs against a pre-quantized DeepSeek-V3 fixture: EP and TP logits match an unsharded reference, and removing the TP half makes that test fail with the scale left whole against a sharded weight. 8-rank smoke runs pass on DeepSeek-V4-Flash (EP) and DeepSeek-V3 (TP); device_map smoke passes on GLM-5.2-NVFP4, DeepSeek-V3, MiniMax-M3-MXFP8 and OLMoE. Also in this pass, from the second review: - `to_local` deprecated rather than deleted; every call site is gone now that `_hf_quantized_needs_local_tp` hands modules local tensors. - `validate_environment` gains its first tests -- the only ones that existed built the frozen FP8 quantizer, which no `from_pretrained` can reach since AUTO_QUANTIZER_MAPPING["fp8"] became FineGrainedHfQuantizer. - `_swizzles_scales` no longer claims block-FP8 never reaches a scaled MMA; it does, via UE8M0, but broadcasts one 128-block scalar in register so there is no per-group grid to lay out. - modelopt remapped to nvfp4 at construction; glob skip-lists moved to the shared `subtree_pattern_to_regex`; GPT-OSS carries swiglu_alpha / swiglu_limit on its config. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…eights reach Every `input_scale` rule lived in the modelopt converters, where the key means NVFP4's second-level global. A plain-FP8 static checkpoint had no rule at all, so Ministral-3 and Mistral-4 would have loaded with the registered default of 1.0 and quantized every activation against the wrong scale, silently. `_static_activation_conversions` routes it to the slot each module holds: one per expert for both MoE layouts, a rename for dense linears. `FineGrainedInputGlobals` serves both meanings now and is named for what it carries; only the NVFP4 gate_up global collapses to one value, because that global is a split of the block scale and a static scale has no block level to absorb an inflated one. Adding the first per-tensor model to the load-path table then reached four defects that nothing else does. A 0-dim scale was sharded as `Shard(ndim - 2)` and had no axis-0 offset to take; the intra-expert rule split an `(E, 1, 1)` scale on a length-1 axis; on-the-fly quantize emitted the dense NVFP4 global as `weight_weight_global_scale` and left the expert input globals on meta, since `ones_like` inherits the device of an unmaterialized slot; and an expert target is a prefix of its own companions', so on save the weight converter claimed `..._input_global_scale` and tried to split one value into gate|up halves. The table and its round-trip assertion are newer than the last run of them: `@slow` kept the suite that covers all of this out of every fast gate. Dropped it — `@require_torch_multi_accelerator` already gates it on the hardware it needs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
| # A 0-dim tensor has no axis to shard, so every rank that owns it takes the whole value | ||
| # (`source[...]`, since a lazy safetensors slice rejects `source[()]`). | ||
| if not source_shape: | ||
| if tensor_idx is not None and not self._owns_expert(tensor_idx, dim_placements): | ||
| return None | ||
| return source[...].to(device=device, dtype=dtype) |
There was a problem hiding this comment.
Its weird we have to do this before, any reason we cannot do this later to enter the "moe" path later down below
| if meta is None: | ||
| return | ||
| placement = Shard(1) if isinstance(module, torch.nn.Embedding) else Shard(meta.ndim - 2) | ||
| # a 0-dim scale is one value for the whole tensor: no axis to split, every rank keeps it |
There was a problem hiding this comment.
an example where this happens, ig some moe weights but interested to see
| def prefers_deepgemm_linear( | ||
| weight: torch.Tensor, | ||
| weight_scale_inv: torch.Tensor, | ||
| *, |
| def _assert_no_post_expert_norm(self: torch.nn.Module) -> None: | ||
| """Mega MoE fuses the routing-weighted reduce into its kernel, so a model's per-expert output | ||
| norm has nowhere to go — it would be dropped silently, so refuse instead. The other DeepGEMM | ||
| arms reduce in `_combine_routed_output` and apply the norm on the rows just before it.""" | ||
| if getattr(self, "has_post_expert_norm", False): | ||
| raise NotImplementedError( | ||
| "DeepGEMM Mega MoE cannot apply this model's per-expert output norm; use " | ||
| "`experts_implementation='deepgemm'`, 'grouped_mm' or 'batched_mm'." | ||
| ) | ||
|
|
||
|
|
||
| def _apply_post_expert_norm(self: torch.nn.Module, rows: torch.Tensor) -> torch.Tensor: | ||
| """The model's per-expert output norm, on the routed rows before the routing weights — the | ||
| same place the reference forwards apply it. A no-op for a model that declares none.""" | ||
| if not getattr(self, "has_post_expert_norm", False): | ||
| return rows | ||
| return self._apply_post_norm(rows) |
|
|
||
| def deepgemm_experts_guards( | ||
| forward=None, | ||
| *, |
There was a problem hiding this comment.
same here, why do we have the * pattern here we shouldnt need it
| # Every companion beside a projection weight — block scales, bias, NVFP4 globals, a | ||
| # static activation scale — is expert-indexed too, so each shards with the weight. The | ||
| # matcher keys on the exact parameter name and falls back only to the owning MODULE, | ||
| # whose entry shards nothing, so without these the weight is this rank's expert slice | ||
| # while its scales are every rank's. gate_up's input global is per-tensor: replicated. |
| # Intra-expert TP: the experts stay whole and the projection's own axis splits, so | ||
| # the per-expert globals, activation scales and the down bias (added after the | ||
| # row-reduce) stay replicated — but the scale grid still has to follow its weight. | ||
| # dim 1 is the output axis and dim 2 the input axis in BOTH the affine and the | ||
| # 5-D swizzled scale container, which is why naming those two suffices. A | ||
| # per-tensor scale has no such grid — one value per expert, `(E, 1, 1)`, whose | ||
| # axes are both 1 — so it joins the replicated companions and only its bias splits. | ||
| blocked = self.quantization_config.weight_block_size is not None | ||
| if projection == "down_proj" and style in ("rowwise", "packed_rowwise"): | ||
| if blocked: | ||
| updated_plan.setdefault(f"{key}_scale_inv", self._expert_shard(2)) | ||
| elif projection != "down_proj" and style in ("packed_colwise", "colwise"): | ||
| # The companions take the WEIGHT's own split. The model declares which one it | ||
| # needs and is right: the interleave into the kernels' row order runs on each | ||
| # rank's shard, AFTER sharding, so the split has to match the layout the | ||
| # CHECKPOINT holds. `packed_colwise` is the strided split a concatenated | ||
| # `[gate; up]` keeps pairs together with; `colwise` the contiguous one a | ||
| # natively interleaved layout (GPT-OSS, `is_concatenated=False`) needs. | ||
| rows = self._expert_shard(1, packed=style == "packed_colwise") | ||
| if blocked: | ||
| updated_plan.setdefault(f"{key}_scale_inv", rows) | ||
| updated_plan.setdefault(f"{key}_bias", rows) |
There was a problem hiding this comment.
shouldnt this be handled with the tp overrides instead? I would expect that tbh
| per_expert = [ | ||
| WeightConverter( | ||
| source_patterns=[ | ||
| r"mlp\.experts\..*\.gate_proj\.weight_scale$", |
There was a problem hiding this comment.
why do we need the .*? we should expect unified layout
| target_patterns=r"\1.activation_scale", | ||
| ) | ||
| ] | ||
| return per_expert + fused + dense |
There was a problem hiding this comment.
Any way we could just do something like in the prev fp8 where we did Fp8Dequantize only on all the scale patterns?
| # "model.layers.1.*" and under `exclude_modules` bare ("self_attn", "layers.0.") | ||
| from ..quantizers.quantizers_utils import subtree_pattern_to_regex | ||
|
|
||
| modules_to_not_convert = [subtree_pattern_to_regex(g) for g in ignore] |
There was a problem hiding this comment.
ah ic but shouldnt we use ours instead, it still feels like its separate PR as it could be a general feat to allow regex
There was a problem hiding this comment.
then we do not need this but modules to convert can just take it as is
Reload: an on-the-fly quantized model lost its swizzle reverse, so saving one wrote swizzled scales a reload could not read back. keep_swizzle_reverse_for_save retains it, and the post-load pass is split into the two delegations it always was, with `covered` computed from source-pattern matches against a fused probe name rather than assumed. TP: a tied lm_head IS the input embedding's parameter, so the embedding_rowwise entry that PretrainedConfig adds for tied weights shards the tensor lm_head reads, while lm_head itself sits outside base_model_tp_plan and gets no style — leaving F.linear with a plain input against a DTensor weight. FSDP resolves the pair; TP had nothing. init_parallel_plans now gives it colwise_rep. It reads the config flag rather than `head.weight is embed.weight` because tying happens later in post_init, where the identity test silently misses. glm4v_moe shipped no expert entries, so its experts were never sharded. Added packed_colwise / rowwise / moe_tp_experts plus shared_experts. DeepseekV4RMSNorm returns the input dtype. V4 pins these norms to FP32 via _keep_in_fp32_modules_strict, so `weight * hidden_states` promoted the result and handed FP32 to compressor projections that ship BF16. The cast-only form is bit-identical where the weight already matches (torch.equal). Tests: pin the fixture to bf16, which exposed three defects (two of them upstream, reproducing with no quantizer at all); add a rounding floor that perturbs every sharded block by one ULP, so the load-path comparison measures against accumulated reduction-order drift instead of an absolute tolerance; size shards against the checkpoint rather than against the other leg. _checkpoint_expert_shape also learns DeepSeek's w1/w3 spelling — it only knew gate_proj/up_proj and the fused layout, so the V4 fp4-experts fixture matched neither and failed outright. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`update_tp_plan` carried `if "Qwen3" in config.__class__.__name__:` and then ASSIGNED a hand-written dense plan over `config.base_model_tp_plan`. Three things were wrong with it: - redundant: the module-keyed entries Qwen3 already ships reach a parameter's scale through the plan matcher's module fallback — `q_proj.weight_scale_inv` resolves to `colwise`, `down_proj.weight_scale_inv` to `rowwise` — so the parameter-keyed duplicates bought nothing; - destructive: assigning dropped what the model ships, including `q_norm`/`k_norm` `replicated_with_grad_allreduce`, and on Qwen3-MoE the whole expert plan (`experts: moe_tp_experts`, `experts.gate_up_proj: packed_colwise`, ...). The companion pass immediately below keys on `moe_tp_experts`, so it then found nothing and the expert scales got no entry either; - over-broad: 25 config classes contain "Qwen3". On the multimodal ones, whose plans live on a sub-config and whose outer config ships none, it stamped a dense text plan onto the outer config where those paths match nothing. Removed. Qwen3 and Qwen3-MoE now keep every shipped entry, Qwen3-MoE additionally gains the three expert companions it was being denied, and a multimodal config's outer plan stays empty. Three assertions cover it, against REAL configs — the existing plan tests build a SimpleNamespace, which a branch keyed on the config CLASS is structurally invisible to. Structure: - the conversion ops move to `integrations/finegrained_conversions.py` (finegrained.py 1631 -> 985). The dependency is one-way; `keep_swizzle_reverse_for_save` goes with them, which is what removes the cycle rather than needing a local import to break it. - `_nvfp4_conversions` splits per checkpoint layout. The two never interact — a checkpoint matches one set and the other never fires — so they were adjacent, not related. - `replace_with_finegrained_layer` hoists the kwargs every swap shares, leaving each branch showing only what is its own. Review comments: - TP styles are registered properly (`moe_experts_colwise` / `moe_experts_packed_colwise` / `moe_experts_rowwise`) instead of `ALL_PARALLEL_STYLES.register` at plan-build time. - the frozen-fp8 deprecation notices fired at IMPORT, for anyone importing `transformers.integrations` at all, one of them through `warnings.warn` rather than the logger. Now one `logger.warning_once` on quantizer construction. - `has_global_scale` dropped: `global_scale_dtype: torch.dtype | None` already answered it. - defensive `getattr` on declared config fields replaced with direct access. - `quantizers_utils.py`, `integrations/tensor_parallel.py` and `integrations/finegrained_fp8.py` are back to their upstream contents and leave this diff entirely. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The ten conversion ops were five self-inverting (reverse is themselves with the
direction flipped) and five paired (reverse names a different class), and every
one of them spelled out its own `__init__` and `reverse_op`. A single
`_FineGrainedOp` base carries the quantizer and defaults `reverse_op` to
`type(self)(..., inverse=not self.inverse)`; the paired ops override it and
ignore the flag. Twelve lines replace fifty-two, and `type(self)` cannot go stale
the way five hardcoded class names could.
- `_moe_operands`: the twelve projection entries were `{kernel_name:
getattr(module, held_name)}` twelve times over. One comprehension states the
rule instead — the kernels always say `gate_up_proj`, the module says
`up_proj` where the model has no gate.
- `FineGrainedExperts.__init__` (104 lines) was reading config and allocating
parameters in one breath. Allocation moves to `_register_projection`, called
once per projection, with the four `_set_optional_parameter` calls folded into
a loop and two loop-invariants hoisted out.
- `_with_expert_layout_ops`: a six-line comment explaining a condition becomes
`sources_the_fused_key`, a named predicate carrying that reasoning; the loop
body goes from fourteen lines to three. The expert-target regex compiles once.
- the `expert_dtype` comment now says why the side-channel exists: DeepSeek-V4 is
MIXED (mxfp4 experts under an "fp8" quant config) and a quant config carries a
single `quant_method` with no way to express it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Two branches decided behaviour by matching a NAME, which is how the Qwen3 plan override stayed invisible — the default arm silently absorbs a miss: - the grouped-linear swap tested `"GroupedLinear" in type(module).__name__`. It now tests `hasattr(module, "n_groups")`, the attribute the branch consumes, so a grouped linear named anything else is still swapped and one named right but shaped wrong no longer reaches an AttributeError. ConvBERT's `GroupedLinearLayer` was never caught (it is an `nn.Module`), so this is hardening rather than a fix. - `_global_role` classified with `"w2" in key`, an unanchored substring for what is a path SEGMENT. Since the default is gate_up, a stray match merged a down global into the wrong bucket silently. Matched on segments now. Readability: - `FineGrainedQuantize` and `FineGrainedDequantize` still carried their own `__init__`; the base-class collapse only matched the form with a default. - `_weight_holder`'s pair was indexed as `holder[0]` / `holder[1]` in four places, and is now unpacked into `module, scale_name`. - `_recontain` -> `_as_container` and `_swizzles_scales` -> `_holds_swizzled_scales`: both were coined, and the codebase already says "container" (`scale_container_dtype`) and "holds" (`holds_interleaved_gate_up`). - `q` / `s` -> `blocked` / `per_block` in the dequant reshape. - each file opens with a short statement of what it is and where its neighbour lives, rather than none at all. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Groups: every quant config is GROUPED internally, and the flat fields a single-format checkpoint ships normalise into one catch-all group. A mixed checkpoint needs more than one -- DeepSeek-V4 is W4A4 mxfp4 EXPERTS over block-FP8 linears -- and compressed-tensors/modelopt already describe that parametrically under `config_groups`, which was being dropped. `group_for(name)` answers with the group that owns a module, so the swap stops asking the config which format the whole model is. A targeted group wins over the catch-all whatever order they were declared in, because `to_json_string` sorts keys and a config that round-tripped through `config.json` has lost that order. `split_out_experts` folds the legacy `config.expert_dtype` side-channel into the two groups it always meant. Four defects, found bisecting a red load-path suite: - the experts' static `activation_scale` was allocated without an explicit dtype, so `from_pretrained` gave it the checkpoint's bf16. Its two siblings kept theirs; a loop collapse made the odd one out invisible. - that scale was never WRITTEN. The loader materialises a key no checkpoint supplies with `torch.empty_like` and `_init_weights` has no branch for a scale, so it reached the kernels as uninitialised memory -- read back as 0 and -0.0244 on a single device. A zero divides the activations by zero, which is where the NaN logits came from, and the randomness is why they moved between runs. `FineGrainedQuantize` now writes every calibration-fed slot as the identity, which is what it already did for the NVFP4 global. - `_quantize_block_fp8` refused a shape the block does not divide. The format's own producers do not: DeepSeek-V3 ships `kv_a_proj_with_mqa` as (576, 7168) E4M3 against a 128x128 block with a (5, 56) scale grid. It pads the trailing partial block now and rounds the grid up to match, so on-the-fly quantization lands what the checkpoint would have. - that refusal was doubling as the filter for tensors that are not weights: an expert bias is 2-D and `E % 128` never divided, so it fell out there. Turned away against `_weight_holder` instead, which already knew. With shapes no longer declined, a full-precision weight in a quantized slot can only mean the checkpoint does not quantize a module it targets -- and the forward has no path for it -- so the three `element_size() > 1` fallbacks go and the check raises rather than warns. Tests: dtype invariance under both default dtypes (rather than asserting fp32, so it catches the next slot too), identity activation scales, ceil-rounded scale grids, and bias pass-through. Each verified to fail without its fix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
- `_quantize_block_fp8` loses both `if pad_m or pad_n` branches: padding by `-rows % block_m` is already a no-op when the shape tiles, and cropping back is then a full slice, which `contiguous` returns unchanged. - `_identity_activation_scales` names the projection prefix once instead of probing `hasattr` per slot -- a dense linear's stem IS `weight`, which is what distinguishes it from the experts' `<proj>_`. - the two new comment blocks say the same thing in half the lines. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`FineGrainedGroupedLinear` took `block_size`/`activation_scheme`/`scale_fmt` but not `weight_format`/`activation_format`, so it silently built its parent's fp8 default -- right for DeepSeek-V4's block-FP8 grouped linears today, wrong the moment a group says otherwise. It takes the pair now, which is also what let the swap drop the comprehension that picked three keys back out of `storage_for`: all three branches pass the same mapping. The experts branch was 27 of the dispatch's 40 lines, so the dispatch itself was the hard part to find. `_quantized_experts` holds the flags-twice-over dance and the post-norm carry-over, and the three-way choice now reads at a glance. `FineGrainedQuantize`'s docstring still said shapes that don't tile pass through, which stopped being true when the block quantizer learned to pad. It names `_weight_holder` as what actually decides, since rank does not. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
CI recapDashboard: View test results in Grafana
|
|
[For maintainers] Suggested jobs to run (before merge) run-slow: deepseek_v4, glm4v_moe, gpt_oss, openai_privacy_filter, finegrained_fp8, mxfp4, nvfp4 |
What does this PR do?
Fixes # (issue)
Code Agent Policy
The Transformers repo is currently being overwhelmed by a large number of PRs and issue comments written by
code agents. These often are low-quality, or fix extremely minor issues that occur rarely or never in practice.
As a result, we're instituting a rule that first-time contributors should not use code agents to submit PRs or issues.
We'd also ask autonomous "OpenClaw"-like agents not to open any PRs or issues.
Issues/PRs from first-time contributors that violate this rule will probably just be closed without review, and we
might block you, especially if you open more than one or appear to be deliberately ignoring this. We especially do not
want new contributors to jump in on random issues to contribute an agent-written fix. This creates lots of noise
for reviewers and other users and will almost certainly get you blocked.
For more information, please read
CONTRIBUTING.md.Before submitting
Pull Request checks?
to it if that's the case.
Who can review?
Anyone in the community is free to review the PR once the tests have passed. Feel free to tag
members/contributors who may be interested in your PR.