Transformers documentation: AWQ (Activation-aware Weight Quantization)
- Type
- other
- Venue
- Hugging Face Transformers docs
- Year
- 2026
- Source
- huggingface
- Access
- free
- Language
- en
- Added
- 2026-08-14T18:56:35Z
- Verified
- 2026-08-14T18:56:35Z
Summary
Living Transformers docs for Activation-aware Weight Quantization: identify AWQ checkpoints via quantization_config.quant_method=awq (typically 4-bit, group_size 128, zero_point, gemm). from_pretrained loads autoawq/llm-awq weights with other tensors in fp16 by default; dtype and device_map control the rest. Notes AutoAWQ pins Transformers to 4.47.1. Optional AwqConfig fused modules (Llama/Mistral out of the box; fuse_max_seq_len + do_fuse) roughly double decode tok/s vs unfused on TheBloke/Mistral-7B-OpenOrca-AWQ at batch=1 (e.g. 32/32: 80.3 vs 38.5 tok/s, 4.00 vs 4.50 GB). ExLlamaV2 via version=exllama (AMD GPUs supported). Cannot combine fused modules with FlashAttention2. Method paper is Lin et al. AWQ (arXiv 2306.00978, MLSys 2024).
Keywords
blog · awq · quantization · transformers · autoawq · exllamav2 · llm-compression
Topics
LLM quantization, inference
Research notes
- Primary: Hugging Face Transformers AWQ docs (posted link). Underlying method paper https://arxiv.org/abs/2306.00978 (Lin, Tang, Tang, Yang, Chen, Wang, Xiao, Dang, Gan, Han; MIT Han Lab; github.com/mit-han-lab/llm-awq). Libraries: autoawq, llm-awq, optimum-intel. Example checkpoints TheBloke/zephyr-7B-alpha-AWQ and TheBloke/Mistral-7B-OpenOrca-AWQ. Demo notebook linked from the page. Documentation, not a new corpus, so no datasets_local row.