← Back to explorer

Transformers documentation: AWQ (Activation-aware Weight Quantization)

Type
other
Venue
Hugging Face Transformers docs
Year
2026
Source
huggingface
Access
free
Language
en
Added
2026-08-14T18:56:35Z
Verified
2026-08-14T18:56:35Z

Summary

Living Transformers docs for Activation-aware Weight Quantization: identify AWQ checkpoints via quantization_config.quant_method=awq (typically 4-bit, group_size 128, zero_point, gemm). from_pretrained loads autoawq/llm-awq weights with other tensors in fp16 by default; dtype and device_map control the rest. Notes AutoAWQ pins Transformers to 4.47.1. Optional AwqConfig fused modules (Llama/Mistral out of the box; fuse_max_seq_len + do_fuse) roughly double decode tok/s vs unfused on TheBloke/Mistral-7B-OpenOrca-AWQ at batch=1 (e.g. 32/32: 80.3 vs 38.5 tok/s, 4.00 vs 4.50 GB). ExLlamaV2 via version=exllama (AMD GPUs supported). Cannot combine fused modules with FlashAttention2. Method paper is Lin et al. AWQ (arXiv 2306.00978, MLSys 2024).

Keywords

blog · awq · quantization · transformers · autoawq · exllamav2 · llm-compression

Topics

LLM quantization, inference

Research notes

  • Primary: Hugging Face Transformers AWQ docs (posted link). Underlying method paper https://arxiv.org/abs/2306.00978 (Lin, Tang, Tang, Yang, Chen, Wang, Xiao, Dang, Gan, Han; MIT Han Lab; github.com/mit-han-lab/llm-awq). Libraries: autoawq, llm-awq, optimum-intel. Example checkpoints TheBloke/zephyr-7B-alpha-AWQ and TheBloke/Mistral-7B-OpenOrca-AWQ. Demo notebook linked from the page. Documentation, not a new corpus, so no datasets_local row.