← Back to explorer

Disaggregated Quantization: Specializing LLM Prefill and Decode

Type
paper
Venue
IST Austria Distributed Algorithms and Systems Lab (arXiv 2026-09-21; HF #2 Paper of the Day)
Year
2026
Source
paper_page
Access
public
Language
en
Added
2026-09-30
Verified
2026-09-30

Summary

Disaggregated quantization (DQ): separate computation formats, weights, and storage placement specialized to prefill vs decode -- low-precision arithmetic accelerates prompt processing, compact weights reduce memory traffic during generation. Removing activation quantization on decode improves accuracy on decode-heavy tasks without added inference cost (Qwen 3, Gemma 3). Separate compute-native prefill weights accelerate prompt processing vs weight-only inference while matching/exceeding its accuracy at 2-3-bit decode. Training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint (Qwen3.8-27B GGUF decoder). "Offloaded disaggregated prefill" (ODP) streams the extra prefill checkpoint from SSD, amortizing loading over prompt length: 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp. Accuracy evaluated under disaggregated serving in vLLM; shared-weight format disaggregation validated via PTQ on models up to 2.8T parameters.

Keywords

quantization · inference · prefill · decode · NVFP4 · low-bit inference · serving

Topics

quantization, inference, prefill, decode, NVFP4, low-bit, vLLM, llama.cpp

Research notes

  • Discovery: https://huggingface.co/papers/2609.26333
  • arXiv: https://arxiv.org/abs/2609.26333
  • Code: https://github.com/IST-DASLab/disaggregated-quantization
  • Prefiller models: https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-NVFP4-prefiller
  • No license shown on paper page; no datasets linked.
  • Connects to the collection's quantization/low-bit inference entries.