Skip to content

Commit 88f7d04

Browse files
xiaoyu-workCopilot
andauthored
Support quantized Gemma 4 per-layer input projections (#762)
Use the decoder quantization factory for PLE gates and projections. Cover GPTQ and Olive packed weights, component layouts, and floating-point behavior. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
1 parent 26a712b commit 88f7d04

1 file changed

Lines changed: 5 additions & 3 deletions

File tree

‎src/mobius/models/gemma4.py‎

Lines changed: 5 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -13,7 +13,8 @@
1313
- Dual head_dim: local sliding-window layers use ``config.head_dim``; global
1414
full-attention layers use ``config.global_head_dim``.
1515
- Dual RoPE: different ``rope_theta`` and ``partial_rotary_factor`` per layer type.
16-
- Per-layer input gating (disabled when ``hidden_size_per_layer_input == 0``).
16+
- Per-layer input gating with decoder-matched quantized projections
17+
(disabled when ``hidden_size_per_layer_input == 0``).
1718
- Vision encoder: pre-patchified input ``[B, N, 3*P^2]`` with 2D position lookup,
1819
bidirectional attention, 4-norm structure, and scale-then-project pooling.
1920
- Vision projector: scale-free RMSNorm -> Linear (matches ``embed_vision`` weights).
@@ -1577,10 +1578,11 @@ def __init__(self, config: Gemma4Config, layer_idx: int):
15771578

15781579
self._per_layer_dim = config.hidden_size_per_layer_input
15791580
if self._per_layer_dim > 0:
1580-
self.per_layer_input_gate = Linear(
1581+
linear_class = _text_linear_class(config) or Linear
1582+
self.per_layer_input_gate = linear_class(
15811583
config.hidden_size, self._per_layer_dim, bias=False
15821584
)
1583-
self.per_layer_projection = Linear(
1585+
self.per_layer_projection = linear_class(
15841586
self._per_layer_dim, config.hidden_size, bias=False
15851587
)
15861588
self.post_per_layer_input_norm = RMSNorm(

0 commit comments

Comments
 (0)