Skip to content

Remove references to torchao's AffineQuantizedTensor - #45321

Open
andrewor14 wants to merge 2 commits into
huggingface:mainfrom
andrewor14:delete-aqt
Open

andrewor14 wants to merge 2 commits into
huggingface:mainfrom
andrewor14:delete-aqt

Conversation

@andrewor14

Copy link
Copy Markdown
Contributor

Summary: TorchAO recently deprecated AffineQuantizedTensor and related classes (pytorch/ao#2752). These will be removed in the next release. We should remove references of these classes in transformers before then.

Test Plan:

python -m pytest -s -v tests/quantization/torchao_integration/test_torchao.py

@andrewor14

Copy link
Copy Markdown
Contributor Author

Hi @SunMarc can you help review this?

@jerryzh168

Copy link
Copy Markdown
Contributor

also cc @MekkCyber

Summary: TorchAO recently deprecated AffineQuantizedTensor and
related classes (pytorch/ao#2752). These will be removed in the
next release. We should remove references of these classes in
transformers before then.

Test Plan:
```
python -m pytest -s -v tests/quantization/torchao_integration/test_torchao.py
```
andrewor14 added a commit to andrewor14/peft that referenced this pull request Apr 8, 2026
**Summary:** TorchAO recently deprecated AffineQuantizedTensor
and related classes (pytorch/ao#2752). These will be removed
in the next release. We should remove references of these
classes in peft before then.

Similar to:
- huggingface/diffusers#13405
- huggingface/transformers#45321

**Test Plan:** CI

@SunMarc SunMarc left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good to me thanks !

@github-actions

github-actions Bot commented Apr 9, 2026

Copy link
Copy Markdown
Contributor

[For maintainers] Suggested jobs to run (before merge)

run-slow: torchao_integration

@HuggingFaceDocBuilderDev

Copy link
Copy Markdown

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

@SunMarc
SunMarc enabled auto-merge April 9, 2026 13:24
@SunMarc
SunMarc added this pull request to the merge queue Apr 9, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Apr 9, 2026
@SunMarc
SunMarc added this pull request to the merge queue Apr 9, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Apr 9, 2026
@SunMarc
SunMarc added this pull request to the merge queue Apr 9, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Apr 9, 2026
guarin pushed a commit to guarin/transformers that referenced this pull request Aug 5, 2026
…7797)

torchao removed the torchao/dtypes package (AffineQuantizedTensor and the
Layout classes). This updates the transformers torchao integration to stop
importing from torchao.dtypes so it no longer crashes against current
torchao, mirroring the approach in transformers huggingface#45321.

Changes:
- integrations/torchao.py: drop the `_quantization_type` helper and its
  `from torchao.dtypes import AffineQuantizedTensor` isinstance branch;
  `_linear_extra_repr` now detects quantized weights inline via
  `isinstance(self.weight, TorchAOBaseTensor)` (from torchao.utils).
- tests/quantization/torchao_integration/test_torchao.py: replace the
  `from torchao.dtypes import AffineQuantizedTensor` import with
  `from torchao.utils import TorchAOBaseTensor` and update the isinstance
  assertions accordingly.
- docs/source/en/quantization/torchao.md: remove the MarlinSparseLayout
  ("int4-weight-only-24sparse") examples and migrate the Int4CPULayout
  example to `Int4WeightOnlyConfig(int4_packing_format=PLAIN_INT32)`.

Test plan:
- `python -c "import transformers.integrations.torchao"` imports cleanly.
- `RUN_SLOW=1 pytest tests/quantization/torchao_integration/test_torchao.py`:
  all 18 TorchAOBaseTensor isinstance / FqnToConfig tests pass. The 3
  remaining failures (test_int4wo_quant, test_int4wo_offload,
  test_serialization_accelerator_*_Int4WeightOnlyConfig) are pre-existing
  generation-output text comparisons pinned to torchao int4 kernel numerics
  on the test hardware, unrelated to this change.

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants