A curated collection of papers, benchmarks, surveys, and tools for model quantization, covering low-bit networks, LLMs, multimodal and generative models, vector and lattice quantization, and efficient deployment.
2,449
336 commits
updated Sep 21, 2026
Awesome Model Quantization is a curated, continuously updated collection of papers, benchmarks, surveys, and open-source implementations on neural network and model quantization. It spans binary and ternary networks, post-training quantization, quantization-aware training, vector and lattice quantization, low-bit LLMs, multimodal and generative models, KV-cache quantization, low-precision training, and hardware-efficient deployment.
Model quantization can be organized along five dimensions:
Methods often combine several dimensions, such as PTQ with rotations and vector codebooks.
Optimization paradigm
Post-Training Quantization (PTQ) converts a pretrained model, often with calibration: GPTQ, SmoothQuant, AWQ, OmniQuant, QuaRot, SpinQuant, FlatQuant, BiLLM. Quantization-Aware Training (QAT) models quantization during optimization: PACT, LSQ, IR-Net. Quantized Fine-Tuning / Parameter-Efficient Fine-Tuning (PEFT) adapts low-bit models: QLoRA, QA-LoRA, LoftQ, IR-QLoRA, L4Q. Data-Free / Zero-Shot Quantization avoids original training data, using model statistics or synthetic samples: ZeroQ, Qimera. Low-Precision Training also reduces precision in training computation or stored states: INT8/FP8 training, 8-bit Optimizers.
Representation / coding structure
Scalar quantization codes individual values; non-uniform, logarithmic, and floating-point quantization change the available levels (AdaLog, LLM-FP4). Vector quantization jointly codes tuples; codebook quantization stores reusable representatives; product / grouped vector quantization partitions vectors into groups (GPTVQ, VPTQ, EPQuant). Lattice quantization uses structured geometric codebooks (QuIP#, NestQuant, grouped lattice vector quantizers). Binary-coded quantization combines binary bases (AnyBCQ); binary / ternary quantization constrains values to two / three levels (IR-Net, BiBERT, PT²-LLM). Mixed precision allocates different bit widths or formats across tensors or groups (HAWQ, SliM-LLM).
Transformation / error handling
Rotation / orthogonal transforms redistribute coordinates (QuaRot, SpinQuant); outlier smoothing / redistribution balances quantization difficulty (SmoothQuant, AWQ). Residual / low-rank reconstruction models remaining errors or outliers (LQER, SVDQuant); error compensation corrects quantization effects (GPTQ, First-Order Error Matters). Saliency-aware / Hessian-aware quantization uses importance or curvature to guide precision, reconstruction, or rounding (HAWQ, GPTQ, BiLLM). These techniques can accompany scalar or structured coding.
Quantized object
Weights (GPTQ, AWQ); activations (PACT); weight + activation (SmoothQuant, BiBERT); KV cache (KIVI, KVQuant, ZipCache, PM-KVQ); training states / optimizer states (ActNN, 8-bit Optimizers); gradients / communication (DoReFa-Net, SDP4Bit). Weight bit width does not imply the same activation, accumulator, or cache precision.
Model family / deployment
CNNs / classical vision (XNOR-Net, BRECQ); Vision Transformers (PTQ4ViT); Large Language Models (GPTQ, QLoRA); multimodal / VLM / VLA (Q-VLM, MQuant, AutoQVLA); diffusion / generative models (PTQ4DM, Q-Diffusion, PTQD, ViDiT-Q, SVDQuant, BinaryDM, Q-VDiT, S²Q-VDiT, QuantSparse); Mamba / state space models (Quamba2, SSDi8); graph / point cloud models (EPQuant, BiPointNet); edge / embedded / hardware-oriented systems (HAQ, FINN, LUT-GEMM).
For the structured-coding lineage, QuIP introduces incoherence processing for low-bit LLM quantization; QuIP# connects this direction to lattice codebooks, while QTIP uses trellis coding. GPTVQ, VPTQ, and NestQuant explore vector or lattice representations. TurboQuant and RaBitQ are also retained for their vector-quantization methodology; RaBitQ targets approximate nearest-neighbor search rather than LLM weight quantization.
Bit width alone does not specify storage overhead, arithmetic precision, or deployment speed.
Selected works grouped by technical approach, with authors and short method summaries. Each paper has one primary home here; the yearly collection includes the wider literature.
Reading the links: Scholar opens a title search on Google Scholar, where citation counts can be viewed. Star badges show the linked GitHub repository’s stars, which may cover several papers. These are discovery aids, not rankings.
From binary weights and learned codebooks to integer arithmetic, learned quantizers and reconstruction-based PTQ.
BinaryConnect: Training Deep Neural Networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, Jean-Pierre David
NeurIPS 2015 · Neural Networks QAT Binary Weights · Paper · Code · Scholar
Trains neural networks with binary weights during forward and backward propagation.
XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, Ali Farhadi
ECCV 2016 · CNN Binary Weight + Activation · Paper · Code · Scholar
Approximates convolutions with binary weights and inputs for efficient CNN inference.
Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Song Han, Huizi Mao, William J. Dally
ICLR 2016 · CNN Weight Sharing Codebook · Paper · Scholar
Combines pruning, trained weight sharing and Huffman coding, connecting learned quantization to compressed model storage.
PACT: Parameterized Clipping Activation for Quantized Neural Networks
Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, Kailash Gopalakrishnan
ICLR 2018 · CNN QAT Activations · Paper · Scholar
Learns activation clipping thresholds to support low-bit network training.
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, Dmitry Kalenichenko
CVPR 2018 · QAT INT8 Integer-Only Inference · Paper · Scholar
Co-designs quantization-aware training and integer arithmetic for mobile inference, including scale and zero-point handling.
HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-Precision
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, Kurt Keutzer
ICCV 2019 · Mixed Precision Hessian-Aware · Paper · Scholar
Uses Hessian information to guide mixed-precision neural network quantization.
Learned Step Size Quantization
Steven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy, Dharmendra S. Modha
ICLR 2020 · QAT Low-Bit · Paper · Scholar
Learns quantizer step sizes alongside network parameters.
Up or Down? Adaptive Rounding for Post-Training Quantization
Markus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos, Tijmen Blankevoort
ICML 2020 · PTQ Rounding · Paper · Scholar
Optimizes rounding decisions when converting pretrained weights to low precision.
HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks
Zhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami, Michael Mahoney, Kurt Keutzer
NeurIPS 2020 · Mixed Precision Hessian-Aware · Paper · Scholar
Develops trace-weighted Hessian sensitivity for mixed-precision allocation.
BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction
Yuhang Li, Ruihao Gong, Xu Tan, Yang Yang, Peng Hu, Qi Zhang, Fengwei Yu, Wei Wang, Shi Gu
ICLR 2021 · CNN PTQ Reconstruction · Paper · Code · Scholar
Uses block reconstruction to reduce post-training quantization error.
QDrop: Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization
Xiuying Wei, Ruihao Gong, Yuhang Li, Xianglong Liu, Fengwei Yu
ICLR 2022 · PTQ Activations Reconstruction · Paper · Code · Scholar
Randomly bypasses activation quantization during reconstruction to improve low-bit generalization beyond the calibration data.
These methods replace access to the original dataset with model statistics or generated samples; they may still require calibration or optimization.
Data-Free Quantization Through Weight Equalization and Bias Correction
Markus Nagel, Mart van Baalen, Tijmen Blankevoort, Max Welling
ICCV 2019 · CNN PTQ Data-Free · Paper · Scholar
Equalizes channel ranges and corrects quantization-induced bias using model parameters and statistics.
ZeroQ: A Novel Zero Shot Quantization Framework
Yaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami, Michael W. Mahoney, Kurt Keutzer
CVPR 2020 · CNN Data-Free Mixed Precision · Paper · Code · Scholar
Synthesizes calibration inputs from batch-normalization statistics to quantize without the original training dataset.
Diversifying Sample Generation for Accurate Data-Free Quantization
Xiangguo Zhang, Haotong Qin, Yifu Ding, Ruihao Gong, Qinghua Yan, Renshuai Tao, Yuhang Li, Fengwei Yu, Xianglong Liu
CVPR 2021 · Oral · CNN Data-Free Synthetic Data PTQ · Paper · Scholar
Relaxes batch-normalization statistic matching and varies layer-wise emphasis to diversify synthetic calibration samples for data-free quantization.
Diverse Sample Generation: Pushing the Limit of Generative Data-Free Quantization
Haotong Qin, Yifu Ding, Xiangguo Zhang, Jiakai Wang, Xianglong Liu, Jiwen Lu
IEEE TPAMI 2023 · CNN Data-Free PTQ + QAT Sample Diversity · Paper · Code · Scholar
Extends the CVPR 2021 DSG method with theoretical analysis and inter-sample decorrelation, improving synthetic-data generation for both PTQ and QAT.
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, Vikas Chandra
ACL Findings 2024 · LLM QAT Data-Free KV Cache · Paper · Scholar
Uses the pretrained model’s generated text for distillation-based QAT of weights, activations and KV caches.
Weight-only compression, weight–activation quantization and QAT address different deployment needs. Sparse outlier handling and rotations offer complementary ways to control error.
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer
NeurIPS 2022 · Transformer INT8 Mixed Precision · Paper · Code · Scholar
Enables 8-bit matrix multiplication at transformer scale while handling outlier features in higher precision.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh
ICLR 2023 · LLM PTQ Weights · Paper · Code · Scholar
Uses approximate second-order information and error compensation for low-bit weight quantization.
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han
ICML 2023 · LLM PTQ Weight + Activation · Paper · Code · Scholar
Redistributes activation outlier difficulty into weights to enable low-precision matrix multiplication.
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, Song Han
MLSys 2024 · LLM PTQ Weights Saliency-Aware · Paper · Code · Scholar
Uses activation information to guide weight quantization for on-device compression and acceleration.
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, Ping Luo
ICLR 2024 · LLM PTQ Calibration · Paper · Code · Scholar
Optimizes clipping and equivalent transformations to calibrate low-bit LLMs.
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li, Pashmina Cameron, Martin Jaggi, Dan Alistarh, Torsten Hoefler, James Hensman
NeurIPS 2024 · LLM PTQ 4-Bit Rotation · Paper · Code · Scholar
Uses rotations to suppress outliers and enable 4-bit inference.
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, Dan Alistarh
ICLR 2024 · LLM PTQ Weights Sparse Outliers · Paper · Code · Scholar
Separates sensitive outlier weights into a sparse higher-precision component while quantizing the remaining weights.
SqueezeLLM: Dense-and-Sparse Quantization
Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W. Mahoney, Kurt Keutzer
ICML 2024 · LLM PTQ Non-uniform Sparse Outliers · Paper · Code · Scholar
Combines sensitivity-weighted non-uniform scalar quantization with a sparse component for outlier weights.
SpinQuant: LLM Quantization with Learned Rotations
Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, Tijmen Blankevoort
ICLR 2025 · LLM PTQ Learned Rotation · Paper · Code · Scholar
Learns rotations to make LLM representations more amenable to quantization.
FlatQuant: Flatness Matters for LLM Quantization
Yuxuan Sun, Ruikang Liu, Haoli Bai, Han Bao, Kang Zhao, Yuening Li, Jiaxin Hu, Xianzhi Yu, Lu Hou, Chun Yuan, Xin Jiang, Wulong Liu, Jun Yao
ICML 2025 · LLM PTQ Transformation · Paper · Code · Scholar
Targets distribution flatness to improve LLM quantization.
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
Mengzhao Chen, Wenqi Shao, Peng Xu, Jiahao Wang, Peng Gao, Kaipeng Zhang, Ping Luo
ACL 2025 · LLM QAT Low-Bit · Paper · Code · Scholar
Trains block parameters first, then quantization parameters end to end, to reduce the cost of LLM QAT.
QLoRA and related methods adapt low-bit bases with low-rank updates; PV-Tuning also optimizes discrete compressed representations.
QLoRA: Efficient Finetuning of Quantized LLMs
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer
NeurIPS 2023 · LLM PEFT 4-Bit · Paper · Code · Scholar
Fine-tunes low-rank adapters through a frozen 4-bit quantized base model.
QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models
Yuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen, Heng Chang, Hengheng Zhang, Zhengsu Chen, Xiaopeng Zhang, Qi Tian
ICLR 2024 · LLM PEFT Quantization-Aware · Paper · Code · Scholar
Combines quantization-aware optimization with low-rank adaptation.
LoftQ: LoRA-Fine-Tuning-aware Quantization for Large Language Models
Yixiao Li, Yifan Yu, Chen Liang, Pengcheng He, Nikos Karampatziakis, Weizhu Chen, Tuo Zhao
ICLR 2024 · LLM PEFT Low-Bit · Paper · Code · Scholar
Aligns quantization with LoRA initialization to reduce the error encountered during adaptation.
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
Haotong Qin, Xudong Ma, Xingyu Zheng, Xiaoyang Li, Yang Zhang, Shouda Liu, Jie Luo, Xianglong Liu, Michele Magno
ICML 2024 · LLM PEFT Information-Aware · Paper · Code · Scholar
Uses information retention to improve low-bit quantization and LoRA adaptation.
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
Vladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev, Konstantin Burlachenko, Kai Yi, Dan Alistarh, Peter Richtarik
NeurIPS 2024 · LLM Quantized Fine-Tuning Discrete Optimization · Paper · Code · Scholar
Alternates continuous and discrete optimization to fine-tune extremely compressed models, including additive-codebook representations.
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
Hyesung Jeon, Yulhwa Kim, Jae-Joon Kim
ACL 2025 · LLM PEFT QAT · Paper · Scholar
Combines parameter-efficient fine-tuning with quantization-aware training.
Binary CNNs and transformers, post-training binarization, and native ternary pretraining have different training costs and arithmetic requirements.
Bi-Real Net: Enhancing the Performance of 1-bit CNNs With Improved Representational Capability and Advanced Training Algorithm
Zechun Liu, Baoyuan Wu, Wenhan Luo, Xin Yang, Wei Liu, Kwang-Ting Cheng
ECCV 2018 · CNN Binary QAT · Paper · Code · Scholar
Connects real-valued intermediate activations through shortcuts to improve information flow in 1-bit CNNs.
Forward and Backward Information Retention for Accurate Binary Neural Networks
Haotong Qin, Ruihao Gong, Xianglong Liu, Mingzhu Shen, Ziran Wei, Fengwei Yu, Jingkuan Song
CVPR 2020 · CNN QAT Binary 1-Bit · Paper · Code · Scholar
Retains information in both forward activations and backward gradients when training binary neural networks.
ReActNet: Towards Precise Binary Neural Network with Generalized Activation Functions
Zechun Liu, Zhiqiang Shen, Marios Savvides, Kwang-Ting Cheng
ECCV 2020 · CNN Binary QAT · Paper · Code · Scholar
Learns activation shifts and reshaping functions to reduce the accuracy gap between binary and real-valued networks.
BiBERT: Accurate Fully Binarized BERT
Haotong Qin, Yifu Ding, Mingyuan Zhang, Qinghua Yan, Aishan Liu, Qingqing Dang, Ziwei Liu, Xianglong Liu
ICLR 2022 · Transformer NLP Binary Weight + Activation · Paper · Code · Scholar
Targets fully binarized BERT, extending binary networks to transformer language models.
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
Wei Huang, Yangdong Liu, Haotong Qin, Ying Li, Shiming Zhang, Xianglong Liu, Michele Magno, Xiaojuan Qi
ICML 2024 · LLM PTQ Binary Extreme Low-Bit · Paper · Code · Scholar
Uses saliency-aware binarization to push pretrained LLM weights into the extreme low-bit regime.
DB-LLM: Accurate Dual-Binarization for Efficient LLMs
Hong Chen, Chengtao Lv, Liang Ding, Haotong Qin, Xiabin Zhou, Yifu Ding, Xuebo Liu, Min Zhang, Jinyang Guo, Xianglong Liu, Dacheng Tao
ACL Findings 2024 · LLM Dual Binarization Extreme Low-Bit · Paper · Scholar
Uses dual binarization to compress LLMs while retaining accuracy.
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, Furu Wei
arXiv 2024 · LLM QAT Ternary Weights 8-Bit Activations · Paper · Scholar
Extends BitNet’s quantization-aware pretraining to ternary weights; this is a training recipe, distinct from post-training binarization.
ARB-LLM: Alternating Refined Binarizations for Large Language Models
Zhiteng Li, Xianglong Yan, Tianao Zhang, Haotong Qin, Dong Xie, Jiang Tian, Zhongchao Shi, Linghe Kong, Yulun Zhang, Xiaokang Yang
ICLR 2025 · LLM Binary Extreme Low-Bit · Paper · Code · Scholar
Refines alternating binarizations for low-bit LLM representation.
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
Jiaqi Zhao, Miao Zhang, Ming Wang, Yuzhang Shang, Kaihao Zhang, Weili Guan, Yaowei Wang, Min Zhang
ACL 2025 · LLM PTQ Extreme Low-Bit · Paper · Code · Scholar
Explores extremely low-bit post-training quantization for LLMs.
PT²-LLM: Post-Training Ternarization for Large Language Models
Xianglong Yan, Chengzhu Bao, Zhiteng Li, Tianao Zhang, Kaicheng Yang, Haotong Qin, Ruobing Xie, Xingwu Sun, Yulun Zhang
ICLR 2026 · LLM PTQ Ternary · Paper · Code · Scholar
Converts pretrained large language models to ternary representations.
From CNN product quantization to LLM additive, lattice and trellis codes. QuIP provides the incoherence-processing precursor; RaBitQ contributes vector-search methodology.
Compressing Deep Convolutional Networks using Vector Quantization
Yunchao Gong, Liu Liu, Ming Yang, Lubomir Bourdev
arXiv 2014 · CNN Vector Quantization Product Quantization · Paper · Scholar
Studies clustering and product quantization of CNN parameters as early approaches to reducing model storage.
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De Sa
NeurIPS 2023 · LLM PTQ 2-Bit Incoherence · Paper · Code · Scholar
Uses incoherence processing for low-bit quantization with guarantees, forming a precursor to the QuIP# lattice-codebook lineage.
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, Christopher De Sa
ICML 2024 · LLM Lattice Codebook Hadamard · Paper · Code · Scholar
Combines Hadamard incoherence processing with lattice codebooks for LLM quantization.
QTIP: Quantization with Trellises and Incoherence Processing
Albert Tseng, Qingyao Sun, David Hou, Christopher De Sa
NeurIPS 2024 · LLM Trellis Coding Incoherence · Paper · Code · Scholar
Combines trellis-based quantization with incoherence processing for compact LLM representation.
GPTVQ: The Blessing of Dimensionality for LLM Quantization
Mart van Baalen, Andrey Kuzmin, Ivan Koryakovskiy, Markus Nagel, Peter Couperus, Cedric Bastoul, Eric Mahurin, Tijmen Blankevoort, Paul Whatmough
arXiv 2024 · LLM Vector Quantization Weights · Paper · Code · Scholar
Exploits joint quantization of multiple weight coordinates rather than coding each weight independently.
VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models
Yifei Liu, Jicheng Wen, Yang Wang, Shengyu Ye, Li Lyna Zhang, Ting Cao, Cheng Li, Mao Yang
EMNLP 2024 · LLM PTQ Vector Quantization Extreme Low-Bit · Paper · Code · Scholar
Uses vector post-training quantization for extremely low-bit LLM compression.
RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search
Jianyang Gao, Cheng Long
SIGMOD 2024 · Vector Quantization Binary Codes Vector Search · Paper · Code · Scholar
Quantizes high-dimensional vectors with a theoretical error bound for approximate nearest-neighbor search.
Extreme Compression of Large Language Models via Additive Quantization
Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar, Artem Babenko, Dan Alistarh
ICML 2024 · LLM PTQ Additive Codebooks 2–3 Bit · Paper · Code · Scholar
Represents weight vectors as sums of learned codewords and jointly optimizes codebooks within transformer blocks.
NestQuant: nested lattice quantization for matrix products and LLMs
Semyon Savkin, Eitan Porat, Or Ordentlich, Yury Polyanskiy
ICML 2025 · LLM Lattice Matrix Products · Paper · Scholar
Uses nested lattice quantization for matrix products and LLMs.
Learning Grouped Lattice Vector Quantizers for Low-Bit Large Language Models
Xi Zhang, Xiaolin Wu, Jiamang Wang, Weisi Lin
NeurIPS 2025 · LLM Grouped Vector Quantization Lattice · Paper · Scholar
Learns grouped lattice vector quantizers for low-bit LLM representation.
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
Gunho Park, Jeongin Bae, Beomseok Kwon, Byeongwook Kim, Se Jung Kwon, Dongsoo Lee
ICLR 2026 · LLM Binary-Coded Mixed Precision Hardware · Paper · Code · Scholar
Develops flexible binary-coded quantization for hardware-efficient multi-precision LLMs.
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
Amir Zandieh, Majid Daliri, Majid Hadian, Vahab Mirrokni
ICLR 2026 · Vector Quantization Online Distortion · Paper · Scholar
Studies online vector quantization with near-optimal distortion rate.
These methods compress inference-time key and value tensors; their bit widths are separate from model weight precision.
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, Xia Hu
ICML 2024 · LLM KV Cache 2-Bit · Paper · Code · Scholar
Uses asymmetric, tuning-free 2-bit quantization to compress key and value caches.
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael Mahoney, Sophia Shao, Kurt Keutzer, Amir Gholami
NeurIPS 2024 · LLM KV Cache Long Context · Paper · Code · Scholar
Targets long-context inference by reducing the memory occupied by the KV cache.
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
Yefei He, Luoming Zhang, Weijia Wu, Jing Liu, Hong Zhou, Bohan Zhuang
NeurIPS 2024 · LLM KV Cache Salient Tokens · Paper · Code · Scholar
Uses salient-token identification to guide accurate and efficient cache quantization.
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
Tengxuan Liu, Shiyao Li, Jiayi Yang, Tianchen Zhao, Feng Zhou, Xiaohui Song, Guohao Dai, Shengen Yan, Huazhong Yang, Yu Wang
ICLR 2026 · LLM KV Cache Mixed Precision · Paper · Code · Scholar
Progressively quantizes KV caches with mixed precision for long chain-of-thought inference.
Early diffusion PTQ addresses denoising-step sensitivity; later work extends to diffusion transformers, low-rank outlier handling and video generation.
Post-training Quantization on Diffusion Models
Yuzhang Shang, Zhihang Yuan, Bin Xie, Bingzhe Wu, Yan Yan
CVPR 2023 · Diffusion PTQ · Paper · Code · Scholar
Adapts post-training quantization to diffusion model inference.
Q-diffusion: Quantizing Diffusion Models
Xiuyu Li, Yijiang Liu, Long Lian, Huanrui Yang, Zhen Dong, Daniel Kang, Shanghang Zhang, Kurt Keutzer
ICCV 2023 · Diffusion PTQ · Paper · Code · Scholar
Quantizes diffusion models to reduce the cost of iterative generation.
PTQD: Accurate Post-Training Quantization for Diffusion Models
Yefei He, Luping Liu, Jing Liu, Weijia Wu, Hong Zhou, Bohan Zhuang
NeurIPS 2023 · Diffusion PTQ Error Handling · Paper · Code · Scholar
Targets accurate diffusion generation through post-training quantization error handling.
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
Tianchen Zhao, Tongcheng Fang, Haofeng Huang, Rui Wan, Widyadewi Soedarmadji, Enshu Liu, Shiyao Li, Zinan Lin, Guohao Dai, Shengen Yan, Huazhong Yang, Xuefei Ning, Yu Wang
ICLR 2025 · Diffusion Transformer Image + Video Low-Bit · Paper · Code · Scholar
Quantizes diffusion transformers for both image and video generation.
SVDQuant: Absorbing Outliers by Low-Rank Component for 4-Bit Diffusion Models
Muyang Li, Yujun Lin, Zhekai Zhang, Tianle Cai, Xiuyu Li, Junxian Guo, Enze Xie, Chenlin Meng, Jun-Yan Zhu, Song Han
ICLR 2025 · Diffusion 4-Bit Low-Rank · Paper · Code · Scholar
Absorbs outliers into a low-rank component to support 4-bit diffusion models.
BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models
Xingyu Zheng, Xianglong Liu, Haotong Qin, Xudong Ma, Mingyuan Zhang, Haojie Hao, Jiakai Wang, Zixiang Zhao, Jinyang Guo, Michele Magno
ICLR 2025 · Diffusion Binary Weights · Paper · Code · Scholar
Binarizes diffusion model weights for efficient generation.
Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers
Weilun Feng, Chuanguang Yang, Haotong Qin, Xiangqi Li, Yu Wang, Zhulin An, Libo Huang, Boyu Diao, Zixiang Zhao, Yongjun Xu, Michele Magno
ICML 2025 · Video Diffusion Quantization Distillation · Paper · Code · Scholar
Combines quantization and distillation for video-generation diffusion transformers.
S²Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
Weilun Feng, Haotong Qin, Chuanguang Yang, Xiangqi Li, Han Yang, Yuqi Li, Zhulin An, Libo Huang, Michele Magno, Yongjun Xu
NeurIPS 2025 · Video Diffusion Quantization Distillation · Paper · Code · Scholar
Uses salient data and sparse-token distillation to improve quantized video diffusion transformers.
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
Weilun Feng, Chuanguang Yang, Haotong Qin, Mingqiang Wu, Yuqi Li, Xiangqi Li, Zhulin An, Libo Huang, Yulun Zhang, Michele Magno, Yongjun Xu
ICLR 2026 · Video Diffusion Quantization Attention Sparsity · Paper · Code · Scholar
Combines model quantization and attention sparsification to compress video diffusion transformers.
Vision-language models and selective state space models introduce quantization sensitivities beyond those of language-only transformers.
Q-VLM: Post-training Quantization for Large Vision-Language Models
Changyuan Wang, Ziwei Wang, Xiuwei Xu, Yansong Tang, Jie Zhou, Jiwen Lu
NeurIPS 2024 · VLM PTQ Cross-Layer Dependency · Paper · Code · Scholar
Uses cross-layer dependencies to guide block partitioning and quantization of vision-language models.
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
Hung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu, Mohamed S. Abdelfattah, Diana Marculescu
ICML 2025 · Mamba State Space Models PTQ W4A8 / W8A8 · Paper · Code · Scholar
Uses channel clustering and state-group quantization to accommodate the sensitivity of Mamba’s selective state-space computations.
Vision methods and deployment systems connect quantizer design to integer kernels, memory movement and hardware costs.
FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip Leong, Magnus Jahre, Kees Vissers
FPGA 2017 · Binary Networks FPGA Inference · Paper · Code · Scholar
Provides a framework for fast, scalable binarized neural network inference on FPGA hardware.
HAQ: Hardware-Aware Automated Quantization with Mixed Precision
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, Song Han
CVPR 2019 · CNN Mixed Precision Hardware-Aware · Paper · Code · Scholar
Automates mixed-precision quantization with hardware deployment costs in view.
BiPointNet: Binary Neural Network for Point Clouds
Haotong Qin, Zhongang Cai, Mingyuan Zhang, Yifu Ding, Haiyu Zhao, Shuai Yi, Xianglong Liu, Hao Su
ICLR 2021 · Point Clouds Binary QAT 1-Bit · Paper · Code · Scholar
Uses entropy-maximizing aggregation and layer-wise scale recovery to address feature homogenization and scale distortion in binary point-cloud networks.
PTQ4ViT: Post-Training Quantization for Vision Transformers with Twin Uniform Quantization
Zhihang Yuan, Chenhao Xue, Yiqi Chen, Qiang Wu, Guangyu Sun
ECCV 2022 · Vision Transformer PTQ · Paper · Code · Scholar
Uses twin uniform quantization to support post-training compression of vision transformers.
QuantSR: Accurate Low-bit Quantization for Efficient Image Super-Resolution
Haotong Qin, Yulun Zhang, Yifu Ding, Yifan Liu, Xianglong Liu, Martin Danelljan, Fisher Yu
NeurIPS 2023 · Super-Resolution QAT 2–4 Bit · Paper · Code · Scholar
Combines a redistribution-driven learnable quantizer with a depth-dynamic architecture for accurate low-bit image super-resolution.
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
Gunho Park, Baeseong Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, Dongsoo Lee
ICLR 2024 · LLM Quantized Matrix Multiplication Lookup Tables · Paper · Scholar
Uses lookup tables for efficient quantized matrix multiplication in generative language models.
Quant-LLM: Accelerating the Serving of Large Language Models via FP6-Centric Algorithm-System Co-Design on Modern GPUs
Haojun Xia, Zhen Zheng, Xiaoxia Wu, Shiyang Chen, Zhewei Yao, Stephen Youn, Arash Bakhtiari, Michael Wyatt, Donglin Zhuang, Zhongzhu Zhou, Olatunji Ruwase, Yuxiong He, Shuaiwen Leon Song
USENIX ATC 2024 · LLM FP6 GPU Kernels · Paper · Code · Scholar
Uses TC-FPx kernels to support non-power-of-two weight formats efficiently on GPUs; the codebase is also known as FP6-LLM.
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han
MLSys 2025 · LLM W4A8KV4 GPU Serving · Paper · Code · Scholar
Co-designs progressive quantization, attention and GPU kernels to turn reduced precision into serving throughput.
Low-bit floating-point and shared-scale formats complement integer quantization. Format design and model calibration are separate choices.
FP8 Formats for Deep Learning
Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, Naveen Mellempudi, Stuart Oberman, Mohammad Shoeybi, Michael Siu, Hao Wu
arXiv 2022 · FP8 E4M3 / E5M2 Training + Inference · Paper · Scholar
Defines complementary FP8 encodings and evaluates their use in neural network training and inference.
Microscaling Data Formats for Deep Learning
Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, Marius Cornea, Eric Dellinger, Kristof Denolf, Stosic Dusan, Venmugil Elango, Maximilian Golub, Alexander Heinecke, Phil James-Roxby, Dharmesh Jani, Gaurav Kolhe, Martin Langhammer, Ada Li, Levi Melnick, Maral Mesmakhosroshahi, Andres Rodriguez, Michael Schulte, Rasoul Shafipour, Lei Shao, Michael Siu, Pradeep Dubey, Paulius Micikevicius, Maxim Naumov, Colin Verrilli, Ralph Wittig, Doug Burger, Eric Chung
arXiv 2023 · MX Formats Block Scaling Training + Inference · Paper · Code · Scholar
Combines shared block scales with narrow element formats to balance numerical range and hardware efficiency.
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
Shih-yang Liu, Zechun Liu, Xijie Huang, Pingcheng Dong, Kwang-Ting Cheng
EMNLP 2023 · LLM PTQ FP4 · Paper · Code · Scholar
Searches exponent configurations and quantization parameters to handle weight and activation range differences in 4-bit floating point.
Quantization can reduce saved activations, optimizer states, gradient communication or training arithmetic; each targets a different part of the training cost.
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, Milan Vojnovic
NeurIPS 2017 · Training Gradients Communication · Paper · Scholar
Uses randomized gradient quantization with convergence guarantees to trade communication bandwidth against estimator variance.
ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed Training
Jianfei Chen, Lianmin Zheng, Zhewei Yao, Dequan Wang, Ion Stoica, Michael Mahoney, Joseph Gonzalez
ICML 2021 · Training Activations 2-Bit · Paper · Code · Scholar
Compresses saved activations to reduce the memory footprint of neural network training.
8-bit Optimizers via Block-wise Quantization
Tim Dettmers, Mike Lewis, Sam Shleifer, Luke Zettlemoyer
ICLR 2022 · Training Optimizer States 8-Bit · Paper · Code · Scholar
Uses block-wise quantization to reduce optimizer-state memory.
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
Jinda Jia, Cong Xie, Hanlin Lu, Daoce Wang, Hao Feng, Chengming Zhang, Baixi Sun, Haibin Lin, Zhi Zhang, Xin Liu, Dingwen Tao
NeurIPS 2024 · LLM Training Communication 4-Bit · Paper · Code · Scholar
Targets 4-bit communication quantization in sharded data-parallel LLM training.
Optimizing Large Language Model Training Using FP4 Quantization
Ruizhe Wang, Yeyun Gong, Xiao Liu, Guoshuai Zhao, Ziyue Yang, Baining Guo, Zhengjun Zha, Peng Cheng
ICML 2025 · LLM Training FP4 Gradient Estimation · Paper · Scholar
Combines differentiable quantization estimation with outlier handling to stabilize FP4 LLM training.
Choose by evaluation scope: deployment reproducibility, binary networks, LLM capabilities or robustness. Expand a resource below for authors, figures and citation details.
| Resource | What it covers |
|---|---|
| MQBench: Towards Reproducible and Deployable Model Quantization Benchmark NeurIPS 2021 Datasets and Benchmarks Code · Scholar | QAT + deployment Compares quantization algorithms under reproducible settings and hardware backend constraints. |
| BiBench: Benchmarking and Analyzing Network Binarization ICML 2023 Code · Scholar | Binary networks Compares binarization methods across tasks, architectures and deployment settings. |
| Evaluating Quantized Large Language Models ICML 2024 Code · Scholar | Weights, activations + KV cache Evaluates 11 model families on basic NLP, emergent abilities, trustworthiness, dialogue and long-context tasks. |
| LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit EMNLP 2024 Industry Track Code · Scholar | LLM toolkit Compares calibration data, method pipelines and quantization configurations; the toolkit is now LightCompress. |
| An empirical study of LLaMA3 quantization: from LLMs to MLLMs Visual Intelligence 2024 Code · Scholar | LLMs + multimodal Examines low-bit behavior across LLaMA3 language and multimodal models. |
| An Empirical Study of Qwen3 Quantization Visual Intelligence 2026 Code · Scholar | Dense + MoE LLMs Studies quantization across Qwen3 model sizes, architectures and reasoning settings. |
| RobustMQ: Benchmarking Robustness of Quantized Models Visual Intelligence 2023 Scholar | Model robustness Tests quantized models beyond clean accuracy, including robustness under input perturbations. |
Yuhang Li, Mingzhu Shen, Jian Ma, Yan Ren, Mingxin Zhao, Qi Zhang, Ruihao Gong, Fengwei Yu, Junjie Yan
@inproceedings{li2021mqbench,
title={MQBench: Towards Reproducible and Deployable Model Quantization Benchmark},
author={Li, Yuhang and Shen, Mingzhu and Ma, Jian and Ren, Yan and Zhao, Mingxin and Zhang, Qi and Gong, Ruihao and Yu, Fengwei and Yan, Junjie},
booktitle={NeurIPS Datasets and Benchmarks},
year={2021}
}
Haotong Qin, Mingyuan Zhang, Yifu Ding, Aoyu Li, Zhongang Cai, Ziwei Liu, Fisher Yu, Xianglong Liu

@inproceedings{qin2023bibench,
title={BiBench: Benchmarking and Analyzing Network Binarization},
author={Qin, Haotong and Zhang, Mingyuan and Ding, Yifu and Li, Aoyu and Cai, Zhongang and Liu, Ziwei and Yu, Fisher and Liu, Xianglong},
booktitle={International Conference on Machine Learning (ICML)},
year={2023}
}
Shiyao Li, Xuefei Ning, Luning Wang, Tengxuan Liu, Xiangsheng Shi, Shengen Yan, Guohao Dai, Huazhong Yang, Yu Wang
@inproceedings{li2024evaluating,
title={Evaluating Quantized Large Language Models},
author={Li, Shiyao and Ning, Xuefei and Wang, Luning and Liu, Tengxuan and Shi, Xiangsheng and Yan, Shengen and Dai, Guohao and Yang, Huazhong and Wang, Yu},
booktitle={International Conference on Machine Learning},
year={2024},
url={https://proceedings.mlr.press/v235/li24bb.html}
}
Ruihao Gong, Yang Yong, Shiqiao Gu, Yushi Huang, Chengtao Lv, Yunchen Zhang, Dacheng Tao, Xianglong Liu

@inproceedings{gong2024llmc,
title={Llmc: Benchmarking large language model quantization with a versatile compression toolkit},
author={Gong, Ruihao and Yong, Yang and Gu, Shiqiao and Huang, Yushi and Lv, Chengtao and Zhang, Yunchen and Tao, Dacheng and Liu, Xianglong},
booktitle={Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track},
pages={132--152},
year={2024}
}
Wei Huang, Xingyu Zheng, Xudong Ma, Haotong Qin, Chengtao Lv, Hong Chen, Jie Luo, Xiaojuan Qi, Xianglong Liu, Michele Magno

@article{huang2024empirical,
title={An empirical study of llama3 quantization: From llms to mllms},
author={Huang, Wei and Zheng, Xingyu and Ma, Xudong and Qin, Haotong and Lv, Chengtao and Chen, Hong and Luo, Jie and Qi, Xiaojuan and Liu, Xianglong and Magno, Michele},
journal={Visual Intelligence},
volume={2},
number={1},
pages={36},
year={2024},
publisher={Springer}
}
Xingyu Zheng, Yuye Li, Haoran Chu, Yue Feng, Xudong Ma, Zining Wang, Jie Luo, Jinyang Guo, Haotong Qin, Michele Magno, Xianglong Liu

@article{zheng2026empirical,
title={An empirical study of Qwen3 quantization},
author={Zheng, Xingyu and Li, Yuye and Chu, Haoran and Feng, Yue and Ma, Xudong and Wang, Zining and Luo, Jie and Guo, Jinyang and Qin, Haotong and Magno, Michele and Liu, Xianglong},
journal={Visual Intelligence},
volume={4},
pages={11},
year={2026},
doi={10.1007/s44267-026-00114-4}
}
Yisong Xiao, Aishan Liu, Tianyuan Zhang, Haotong Qin, Jinyang Guo, Xianglong Liu

@article{xiao2023robustmq,
title={Robustmq: benchmarking robustness of quantized models},
author={Xiao, Yisong and Liu, Aishan and Zhang, Tianyuan and Qin, Haotong and Guo, Jinyang and Liu, Xianglong},
journal={Visual Intelligence},
volume={1},
number={1},
pages={30},
year={2023},
publisher={Springer}
}
Start with the white paper for practical PTQ/QAT, then choose a survey for broader context or a specific model family. Figures and citation details are available below.
| Resource | What it covers |
|---|---|
| A White Paper on Neural Network Quantization arXiv 2021 Scholar | Practical PTQ + QAT Explains quantizer design, common failure modes and practical post-training and quantization-aware training workflows. |
| A Survey of Quantization Methods for Efficient Neural Network Inference arXiv 2021 Scholar | Foundations + taxonomy Reviews quantization design choices, mixed precision and the trade-offs between model accuracy and efficient inference. |
| Binary Neural Networks: A Survey Pattern Recognition 2020 Scholar | Binary networks Surveys binary network representations, training methods and applications. |
| A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms Neural Networks 2025 Scholar | LLM algorithms + systems Connects low-bit LLM algorithms with numerical formats and inference systems. |
| Low-bit Model Quantization for Deep Neural Networks: A Survey arXiv 2025 Scholar | Broad low-bit methods Maps low-bit quantization methods across neural network architectures and applications. |
Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, Tijmen Blankevoort
@article{nagel2021white,
title={A White Paper on Neural Network Quantization},
author={Nagel, Markus and Fournarakis, Marios and Amjad, Rana Ali and Bondarenko, Yelysei and van Baalen, Mart and Blankevoort, Tijmen},
journal={arXiv preprint arXiv:2106.08295},
year={2021}
}
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, Kurt Keutzer
@article{gholami2021survey,
title={A Survey of Quantization Methods for Efficient Neural Network Inference},
author={Gholami, Amir and Kim, Sehoon and Dong, Zhen and Yao, Zhewei and Mahoney, Michael W. and Keutzer, Kurt},
journal={arXiv preprint arXiv:2103.13630},
year={2021}
}
Haotong Qin, Ruihao Gong, Xianglong Liu, Xiao Bai, Jingkuan Song, Nicu Sebe

@article{Qin:pr20_bnn_survey,
title = "Binary neural networks: A survey",
author = "Haotong Qin and Ruihao Gong and Xianglong Liu and Xiao Bai and Jingkuan Song and Nicu Sebe",
journal = "Pattern Recognition",
volume = "105",
pages = "107281",
year = "2020"
}
Ruihao Gong, Yifu Ding, Zining Wang, Chengtao Lv, Xingyu Zheng, Jinyang Du, Yang Yong, Shiqiao Gu, Haotong Qin, Jinyang Guo, Dahua Lin, Michele Magno, Xianglong Liu

@article{gong2025survey,
title={A survey of low-bit large language models: Basics, systems, and algorithms},
author={Gong, Ruihao and Ding, Yifu and Wang, Zining and Lv, Chengtao and Zheng, Xingyu and Du, Jinyang and Yong, Yang and Gu, Shiqiao and Qin, Haotong and Guo, Jinyang and Lin, Dahua and Magno, Michele and Liu, Xianglong},
journal={Neural Networks},
pages={107856},
year={2025}
}
Kai Liu, Qian Zheng, Kaiwen Tao, Zhiteng Li, Haotong Qin, Wenbo Li, Yong Guo, Xianglong Liu, Linghe Kong, Guihai Chen, Yulun Zhang, Xiaokang Yang

@article{liu2025low,
title={Low-bit Model Quantization for Deep Neural Networks: A Survey},
author={Liu, Kai and Zheng, Qian and Tao, Kaiwen and Li, Zhiteng and Qin, Haotong and Li, Wenbo and Guo, Yong and Liu, Xianglong and Kong, Linghe and Chen, Guihai and Zhang, Yulun and Yang, Xiaokang},
journal={arXiv preprint arXiv:2505.05530},
year={2025}
}
All paper titles and links are kept in this README. Published work is grouped by venue year where verified; otherwise the recorded preprint year is used. Representative works, benchmarks and surveys also appear here for chronological browsing. Within each year, entries are grouped by conference or journal, with preprints at the end.
Truncated — view the full README on GitHub.
(top 30 of 34)
A curated collection of papers, benchmarks, surveys, and tools for model quantization, covering low-bit networks, LLMs, multimodal and generative models, vector and lattice quantization, and efficient deployment.
2,449
336 commits
updated Sep 21, 2026
Awesome Model Quantization is a curated, continuously updated collection of papers, benchmarks, surveys, and open-source implementations on neural network and model quantization. It spans binary and ternary networks, post-training quantization, quantization-aware training, vector and lattice quantization, low-bit LLMs, multimodal and generative models, KV-cache quantization, low-precision training, and hardware-efficient deployment.
Model quantization can be organized along five dimensions:
Methods often combine several dimensions, such as PTQ with rotations and vector codebooks.
Optimization paradigm
Post-Training Quantization (PTQ) converts a pretrained model, often with calibration: GPTQ, SmoothQuant, AWQ, OmniQuant, QuaRot, SpinQuant, FlatQuant, BiLLM. Quantization-Aware Training (QAT) models quantization during optimization: PACT, LSQ, IR-Net. Quantized Fine-Tuning / Parameter-Efficient Fine-Tuning (PEFT) adapts low-bit models: QLoRA, QA-LoRA, LoftQ, IR-QLoRA, L4Q. Data-Free / Zero-Shot Quantization avoids original training data, using model statistics or synthetic samples: ZeroQ, Qimera. Low-Precision Training also reduces precision in training computation or stored states: INT8/FP8 training, 8-bit Optimizers.
Representation / coding structure
Scalar quantization codes individual values; non-uniform, logarithmic, and floating-point quantization change the available levels (AdaLog, LLM-FP4). Vector quantization jointly codes tuples; codebook quantization stores reusable representatives; product / grouped vector quantization partitions vectors into groups (GPTVQ, VPTQ, EPQuant). Lattice quantization uses structured geometric codebooks (QuIP#, NestQuant, grouped lattice vector quantizers). Binary-coded quantization combines binary bases (AnyBCQ); binary / ternary quantization constrains values to two / three levels (IR-Net, BiBERT, PT²-LLM). Mixed precision allocates different bit widths or formats across tensors or groups (HAWQ, SliM-LLM).
Transformation / error handling
Rotation / orthogonal transforms redistribute coordinates (QuaRot, SpinQuant); outlier smoothing / redistribution balances quantization difficulty (SmoothQuant, AWQ). Residual / low-rank reconstruction models remaining errors or outliers (LQER, SVDQuant); error compensation corrects quantization effects (GPTQ, First-Order Error Matters). Saliency-aware / Hessian-aware quantization uses importance or curvature to guide precision, reconstruction, or rounding (HAWQ, GPTQ, BiLLM). These techniques can accompany scalar or structured coding.
Quantized object
Weights (GPTQ, AWQ); activations (PACT); weight + activation (SmoothQuant, BiBERT); KV cache (KIVI, KVQuant, ZipCache, PM-KVQ); training states / optimizer states (ActNN, 8-bit Optimizers); gradients / communication (DoReFa-Net, SDP4Bit). Weight bit width does not imply the same activation, accumulator, or cache precision.
Model family / deployment
CNNs / classical vision (XNOR-Net, BRECQ); Vision Transformers (PTQ4ViT); Large Language Models (GPTQ, QLoRA); multimodal / VLM / VLA (Q-VLM, MQuant, AutoQVLA); diffusion / generative models (PTQ4DM, Q-Diffusion, PTQD, ViDiT-Q, SVDQuant, BinaryDM, Q-VDiT, S²Q-VDiT, QuantSparse); Mamba / state space models (Quamba2, SSDi8); graph / point cloud models (EPQuant, BiPointNet); edge / embedded / hardware-oriented systems (HAQ, FINN, LUT-GEMM).
For the structured-coding lineage, QuIP introduces incoherence processing for low-bit LLM quantization; QuIP# connects this direction to lattice codebooks, while QTIP uses trellis coding. GPTVQ, VPTQ, and NestQuant explore vector or lattice representations. TurboQuant and RaBitQ are also retained for their vector-quantization methodology; RaBitQ targets approximate nearest-neighbor search rather than LLM weight quantization.
Bit width alone does not specify storage overhead, arithmetic precision, or deployment speed.
Selected works grouped by technical approach, with authors and short method summaries. Each paper has one primary home here; the yearly collection includes the wider literature.
Reading the links: Scholar opens a title search on Google Scholar, where citation counts can be viewed. Star badges show the linked GitHub repository’s stars, which may cover several papers. These are discovery aids, not rankings.
From binary weights and learned codebooks to integer arithmetic, learned quantizers and reconstruction-based PTQ.
BinaryConnect: Training Deep Neural Networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, Jean-Pierre David
NeurIPS 2015 · Neural Networks QAT Binary Weights · Paper · Code · Scholar
Trains neural networks with binary weights during forward and backward propagation.
XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, Ali Farhadi
ECCV 2016 · CNN Binary Weight + Activation · Paper · Code · Scholar
Approximates convolutions with binary weights and inputs for efficient CNN inference.
Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
Song Han, Huizi Mao, William J. Dally
ICLR 2016 · CNN Weight Sharing Codebook · Paper · Scholar
Combines pruning, trained weight sharing and Huffman coding, connecting learned quantization to compressed model storage.
PACT: Parameterized Clipping Activation for Quantized Neural Networks
Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, Kailash Gopalakrishnan
ICLR 2018 · CNN QAT Activations · Paper · Scholar
Learns activation clipping thresholds to support low-bit network training.
Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew Howard, Hartwig Adam, Dmitry Kalenichenko
CVPR 2018 · QAT INT8 Integer-Only Inference · Paper · Scholar
Co-designs quantization-aware training and integer arithmetic for mobile inference, including scale and zero-point handling.
HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-Precision
Zhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney, Kurt Keutzer
ICCV 2019 · Mixed Precision Hessian-Aware · Paper · Scholar
Uses Hessian information to guide mixed-precision neural network quantization.
Learned Step Size Quantization
Steven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy, Dharmendra S. Modha
ICLR 2020 · QAT Low-Bit · Paper · Scholar
Learns quantizer step sizes alongside network parameters.
Up or Down? Adaptive Rounding for Post-Training Quantization
Markus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos, Tijmen Blankevoort
ICML 2020 · PTQ Rounding · Paper · Scholar
Optimizes rounding decisions when converting pretrained weights to low precision.
HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks
Zhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami, Michael Mahoney, Kurt Keutzer
NeurIPS 2020 · Mixed Precision Hessian-Aware · Paper · Scholar
Develops trace-weighted Hessian sensitivity for mixed-precision allocation.
BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction
Yuhang Li, Ruihao Gong, Xu Tan, Yang Yang, Peng Hu, Qi Zhang, Fengwei Yu, Wei Wang, Shi Gu
ICLR 2021 · CNN PTQ Reconstruction · Paper · Code · Scholar
Uses block reconstruction to reduce post-training quantization error.
QDrop: Randomly Dropping Quantization for Extremely Low-bit Post-Training Quantization
Xiuying Wei, Ruihao Gong, Yuhang Li, Xianglong Liu, Fengwei Yu
ICLR 2022 · PTQ Activations Reconstruction · Paper · Code · Scholar
Randomly bypasses activation quantization during reconstruction to improve low-bit generalization beyond the calibration data.
These methods replace access to the original dataset with model statistics or generated samples; they may still require calibration or optimization.
Data-Free Quantization Through Weight Equalization and Bias Correction
Markus Nagel, Mart van Baalen, Tijmen Blankevoort, Max Welling
ICCV 2019 · CNN PTQ Data-Free · Paper · Scholar
Equalizes channel ranges and corrects quantization-induced bias using model parameters and statistics.
ZeroQ: A Novel Zero Shot Quantization Framework
Yaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami, Michael W. Mahoney, Kurt Keutzer
CVPR 2020 · CNN Data-Free Mixed Precision · Paper · Code · Scholar
Synthesizes calibration inputs from batch-normalization statistics to quantize without the original training dataset.
Diversifying Sample Generation for Accurate Data-Free Quantization
Xiangguo Zhang, Haotong Qin, Yifu Ding, Ruihao Gong, Qinghua Yan, Renshuai Tao, Yuhang Li, Fengwei Yu, Xianglong Liu
CVPR 2021 · Oral · CNN Data-Free Synthetic Data PTQ · Paper · Scholar
Relaxes batch-normalization statistic matching and varies layer-wise emphasis to diversify synthetic calibration samples for data-free quantization.
Diverse Sample Generation: Pushing the Limit of Generative Data-Free Quantization
Haotong Qin, Yifu Ding, Xiangguo Zhang, Jiakai Wang, Xianglong Liu, Jiwen Lu
IEEE TPAMI 2023 · CNN Data-Free PTQ + QAT Sample Diversity · Paper · Code · Scholar
Extends the CVPR 2021 DSG method with theoretical analysis and inter-sample decorrelation, improving synthetic-data generation for both PTQ and QAT.
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Zechun Liu, Barlas Oguz, Changsheng Zhao, Ernie Chang, Pierre Stock, Yashar Mehdad, Yangyang Shi, Raghuraman Krishnamoorthi, Vikas Chandra
ACL Findings 2024 · LLM QAT Data-Free KV Cache · Paper · Scholar
Uses the pretrained model’s generated text for distillation-based QAT of weights, activations and KV caches.
Weight-only compression, weight–activation quantization and QAT address different deployment needs. Sparse outlier handling and rotations offer complementary ways to control error.
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Tim Dettmers, Mike Lewis, Younes Belkada, Luke Zettlemoyer
NeurIPS 2022 · Transformer INT8 Mixed Precision · Paper · Code · Scholar
Enables 8-bit matrix multiplication at transformer scale while handling outlier features in higher precision.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, Dan Alistarh
ICLR 2023 · LLM PTQ Weights · Paper · Code · Scholar
Uses approximate second-order information and error compensation for low-bit weight quantization.
SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language Models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, Song Han
ICML 2023 · LLM PTQ Weight + Activation · Paper · Code · Scholar
Redistributes activation outlier difficulty into weights to enable low-precision matrix multiplication.
AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, Song Han
MLSys 2024 · LLM PTQ Weights Saliency-Aware · Paper · Code · Scholar
Uses activation information to guide weight quantization for on-device compression and acceleration.
OmniQuant: Omnidirectionally Calibrated Quantization for Large Language Models
Wenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu, Lirui Zhao, Zhiqian Li, Kaipeng Zhang, Peng Gao, Yu Qiao, Ping Luo
ICLR 2024 · LLM PTQ Calibration · Paper · Code · Scholar
Optimizes clipping and equivalent transformations to calibrate low-bit LLMs.
QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li, Pashmina Cameron, Martin Jaggi, Dan Alistarh, Torsten Hoefler, James Hensman
NeurIPS 2024 · LLM PTQ 4-Bit Rotation · Paper · Code · Scholar
Uses rotations to suppress outliers and enable 4-bit inference.
SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, Dan Alistarh
ICLR 2024 · LLM PTQ Weights Sparse Outliers · Paper · Code · Scholar
Separates sensitive outlier weights into a sparse higher-precision component while quantizing the remaining weights.
SqueezeLLM: Dense-and-Sparse Quantization
Sehoon Kim, Coleman Hooper, Amir Gholami, Zhen Dong, Xiuyu Li, Sheng Shen, Michael W. Mahoney, Kurt Keutzer
ICML 2024 · LLM PTQ Non-uniform Sparse Outliers · Paper · Code · Scholar
Combines sensitivity-weighted non-uniform scalar quantization with a sparse component for outlier weights.
SpinQuant: LLM Quantization with Learned Rotations
Zechun Liu, Changsheng Zhao, Igor Fedorov, Bilge Soran, Dhruv Choudhary, Raghuraman Krishnamoorthi, Vikas Chandra, Yuandong Tian, Tijmen Blankevoort
ICLR 2025 · LLM PTQ Learned Rotation · Paper · Code · Scholar
Learns rotations to make LLM representations more amenable to quantization.
FlatQuant: Flatness Matters for LLM Quantization
Yuxuan Sun, Ruikang Liu, Haoli Bai, Han Bao, Kang Zhao, Yuening Li, Jiaxin Hu, Xianzhi Yu, Lu Hou, Chun Yuan, Xin Jiang, Wulong Liu, Jun Yao
ICML 2025 · LLM PTQ Transformation · Paper · Code · Scholar
Targets distribution flatness to improve LLM quantization.
EfficientQAT: Efficient Quantization-Aware Training for Large Language Models
Mengzhao Chen, Wenqi Shao, Peng Xu, Jiahao Wang, Peng Gao, Kaipeng Zhang, Ping Luo
ACL 2025 · LLM QAT Low-Bit · Paper · Code · Scholar
Trains block parameters first, then quantization parameters end to end, to reduce the cost of LLM QAT.
QLoRA and related methods adapt low-bit bases with low-rank updates; PV-Tuning also optimizes discrete compressed representations.
QLoRA: Efficient Finetuning of Quantized LLMs
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke Zettlemoyer
NeurIPS 2023 · LLM PEFT 4-Bit · Paper · Code · Scholar
Fine-tunes low-rank adapters through a frozen 4-bit quantized base model.
QA-LoRA: Quantization-Aware Low-Rank Adaptation of Large Language Models
Yuhui Xu, Lingxi Xie, Xiaotao Gu, Xin Chen, Heng Chang, Hengheng Zhang, Zhengsu Chen, Xiaopeng Zhang, Qi Tian
ICLR 2024 · LLM PEFT Quantization-Aware · Paper · Code · Scholar
Combines quantization-aware optimization with low-rank adaptation.
LoftQ: LoRA-Fine-Tuning-aware Quantization for Large Language Models
Yixiao Li, Yifan Yu, Chen Liang, Pengcheng He, Nikos Karampatziakis, Weizhu Chen, Tuo Zhao
ICLR 2024 · LLM PEFT Low-Bit · Paper · Code · Scholar
Aligns quantization with LoRA initialization to reduce the error encountered during adaptation.
Accurate LoRA-Finetuning Quantization of LLMs via Information Retention
Haotong Qin, Xudong Ma, Xingyu Zheng, Xiaoyang Li, Yang Zhang, Shouda Liu, Jie Luo, Xianglong Liu, Michele Magno
ICML 2024 · LLM PEFT Information-Aware · Paper · Code · Scholar
Uses information retention to improve low-bit quantization and LoRA adaptation.
PV-Tuning: Beyond Straight-Through Estimation for Extreme LLM Compression
Vladimir Malinovskii, Denis Mazur, Ivan Ilin, Denis Kuznedelev, Konstantin Burlachenko, Kai Yi, Dan Alistarh, Peter Richtarik
NeurIPS 2024 · LLM Quantized Fine-Tuning Discrete Optimization · Paper · Code · Scholar
Alternates continuous and discrete optimization to fine-tune extremely compressed models, including additive-codebook representations.
L4Q: Parameter Efficient Quantization-Aware Fine-Tuning on Large Language Models
Hyesung Jeon, Yulhwa Kim, Jae-Joon Kim
ACL 2025 · LLM PEFT QAT · Paper · Scholar
Combines parameter-efficient fine-tuning with quantization-aware training.
Binary CNNs and transformers, post-training binarization, and native ternary pretraining have different training costs and arithmetic requirements.
Bi-Real Net: Enhancing the Performance of 1-bit CNNs With Improved Representational Capability and Advanced Training Algorithm
Zechun Liu, Baoyuan Wu, Wenhan Luo, Xin Yang, Wei Liu, Kwang-Ting Cheng
ECCV 2018 · CNN Binary QAT · Paper · Code · Scholar
Connects real-valued intermediate activations through shortcuts to improve information flow in 1-bit CNNs.
Forward and Backward Information Retention for Accurate Binary Neural Networks
Haotong Qin, Ruihao Gong, Xianglong Liu, Mingzhu Shen, Ziran Wei, Fengwei Yu, Jingkuan Song
CVPR 2020 · CNN QAT Binary 1-Bit · Paper · Code · Scholar
Retains information in both forward activations and backward gradients when training binary neural networks.
ReActNet: Towards Precise Binary Neural Network with Generalized Activation Functions
Zechun Liu, Zhiqiang Shen, Marios Savvides, Kwang-Ting Cheng
ECCV 2020 · CNN Binary QAT · Paper · Code · Scholar
Learns activation shifts and reshaping functions to reduce the accuracy gap between binary and real-valued networks.
BiBERT: Accurate Fully Binarized BERT
Haotong Qin, Yifu Ding, Mingyuan Zhang, Qinghua Yan, Aishan Liu, Qingqing Dang, Ziwei Liu, Xianglong Liu
ICLR 2022 · Transformer NLP Binary Weight + Activation · Paper · Code · Scholar
Targets fully binarized BERT, extending binary networks to transformer language models.
BiLLM: Pushing the Limit of Post-Training Quantization for LLMs
Wei Huang, Yangdong Liu, Haotong Qin, Ying Li, Shiming Zhang, Xianglong Liu, Michele Magno, Xiaojuan Qi
ICML 2024 · LLM PTQ Binary Extreme Low-Bit · Paper · Code · Scholar
Uses saliency-aware binarization to push pretrained LLM weights into the extreme low-bit regime.
DB-LLM: Accurate Dual-Binarization for Efficient LLMs
Hong Chen, Chengtao Lv, Liang Ding, Haotong Qin, Xiabin Zhou, Yifu Ding, Xuebo Liu, Min Zhang, Jinyang Guo, Xianglong Liu, Dacheng Tao
ACL Findings 2024 · LLM Dual Binarization Extreme Low-Bit · Paper · Scholar
Uses dual binarization to compress LLMs while retaining accuracy.
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Shuming Ma, Hongyu Wang, Lingxiao Ma, Lei Wang, Wenhui Wang, Shaohan Huang, Li Dong, Ruiping Wang, Jilong Xue, Furu Wei
arXiv 2024 · LLM QAT Ternary Weights 8-Bit Activations · Paper · Scholar
Extends BitNet’s quantization-aware pretraining to ternary weights; this is a training recipe, distinct from post-training binarization.
ARB-LLM: Alternating Refined Binarizations for Large Language Models
Zhiteng Li, Xianglong Yan, Tianao Zhang, Haotong Qin, Dong Xie, Jiang Tian, Zhongchao Shi, Linghe Kong, Yulun Zhang, Xiaokang Yang
ICLR 2025 · LLM Binary Extreme Low-Bit · Paper · Code · Scholar
Refines alternating binarizations for low-bit LLM representation.
PTQ1.61: Push the Real Limit of Extremely Low-Bit Post-Training Quantization Methods for Large Language Models
Jiaqi Zhao, Miao Zhang, Ming Wang, Yuzhang Shang, Kaihao Zhang, Weili Guan, Yaowei Wang, Min Zhang
ACL 2025 · LLM PTQ Extreme Low-Bit · Paper · Code · Scholar
Explores extremely low-bit post-training quantization for LLMs.
PT²-LLM: Post-Training Ternarization for Large Language Models
Xianglong Yan, Chengzhu Bao, Zhiteng Li, Tianao Zhang, Kaicheng Yang, Haotong Qin, Ruobing Xie, Xingwu Sun, Yulun Zhang
ICLR 2026 · LLM PTQ Ternary · Paper · Code · Scholar
Converts pretrained large language models to ternary representations.
From CNN product quantization to LLM additive, lattice and trellis codes. QuIP provides the incoherence-processing precursor; RaBitQ contributes vector-search methodology.
Compressing Deep Convolutional Networks using Vector Quantization
Yunchao Gong, Liu Liu, Ming Yang, Lubomir Bourdev
arXiv 2014 · CNN Vector Quantization Product Quantization · Paper · Scholar
Studies clustering and product quantization of CNN parameters as early approaches to reducing model storage.
QuIP: 2-Bit Quantization of Large Language Models With Guarantees
Jerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De Sa
NeurIPS 2023 · LLM PTQ 2-Bit Incoherence · Paper · Code · Scholar
Uses incoherence processing for low-bit quantization with guarantees, forming a precursor to the QuIP# lattice-codebook lineage.
QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks
Albert Tseng, Jerry Chee, Qingyao Sun, Volodymyr Kuleshov, Christopher De Sa
ICML 2024 · LLM Lattice Codebook Hadamard · Paper · Code · Scholar
Combines Hadamard incoherence processing with lattice codebooks for LLM quantization.
QTIP: Quantization with Trellises and Incoherence Processing
Albert Tseng, Qingyao Sun, David Hou, Christopher De Sa
NeurIPS 2024 · LLM Trellis Coding Incoherence · Paper · Code · Scholar
Combines trellis-based quantization with incoherence processing for compact LLM representation.
GPTVQ: The Blessing of Dimensionality for LLM Quantization
Mart van Baalen, Andrey Kuzmin, Ivan Koryakovskiy, Markus Nagel, Peter Couperus, Cedric Bastoul, Eric Mahurin, Tijmen Blankevoort, Paul Whatmough
arXiv 2024 · LLM Vector Quantization Weights · Paper · Code · Scholar
Exploits joint quantization of multiple weight coordinates rather than coding each weight independently.
VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models
Yifei Liu, Jicheng Wen, Yang Wang, Shengyu Ye, Li Lyna Zhang, Ting Cao, Cheng Li, Mao Yang
EMNLP 2024 · LLM PTQ Vector Quantization Extreme Low-Bit · Paper · Code · Scholar
Uses vector post-training quantization for extremely low-bit LLM compression.
RaBitQ: Quantizing High-Dimensional Vectors with a Theoretical Error Bound for Approximate Nearest Neighbor Search
Jianyang Gao, Cheng Long
SIGMOD 2024 · Vector Quantization Binary Codes Vector Search · Paper · Code · Scholar
Quantizes high-dimensional vectors with a theoretical error bound for approximate nearest-neighbor search.
Extreme Compression of Large Language Models via Additive Quantization
Vage Egiazarian, Andrei Panferov, Denis Kuznedelev, Elias Frantar, Artem Babenko, Dan Alistarh
ICML 2024 · LLM PTQ Additive Codebooks 2–3 Bit · Paper · Code · Scholar
Represents weight vectors as sums of learned codewords and jointly optimizes codebooks within transformer blocks.
NestQuant: nested lattice quantization for matrix products and LLMs
Semyon Savkin, Eitan Porat, Or Ordentlich, Yury Polyanskiy
ICML 2025 · LLM Lattice Matrix Products · Paper · Scholar
Uses nested lattice quantization for matrix products and LLMs.
Learning Grouped Lattice Vector Quantizers for Low-Bit Large Language Models
Xi Zhang, Xiaolin Wu, Jiamang Wang, Weisi Lin
NeurIPS 2025 · LLM Grouped Vector Quantization Lattice · Paper · Scholar
Learns grouped lattice vector quantizers for low-bit LLM representation.
AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
Gunho Park, Jeongin Bae, Beomseok Kwon, Byeongwook Kim, Se Jung Kwon, Dongsoo Lee
ICLR 2026 · LLM Binary-Coded Mixed Precision Hardware · Paper · Code · Scholar
Develops flexible binary-coded quantization for hardware-efficient multi-precision LLMs.
TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate
Amir Zandieh, Majid Daliri, Majid Hadian, Vahab Mirrokni
ICLR 2026 · Vector Quantization Online Distortion · Paper · Scholar
Studies online vector quantization with near-optimal distortion rate.
These methods compress inference-time key and value tensors; their bit widths are separate from model weight precision.
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, Xia Hu
ICML 2024 · LLM KV Cache 2-Bit · Paper · Code · Scholar
Uses asymmetric, tuning-free 2-bit quantization to compress key and value caches.
KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael Mahoney, Sophia Shao, Kurt Keutzer, Amir Gholami
NeurIPS 2024 · LLM KV Cache Long Context · Paper · Code · Scholar
Targets long-context inference by reducing the memory occupied by the KV cache.
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification
Yefei He, Luoming Zhang, Weijia Wu, Jing Liu, Hong Zhou, Bohan Zhuang
NeurIPS 2024 · LLM KV Cache Salient Tokens · Paper · Code · Scholar
Uses salient-token identification to guide accurate and efficient cache quantization.
PM-KVQ: Progressive Mixed-precision KV Cache Quantization for Long-CoT LLMs
Tengxuan Liu, Shiyao Li, Jiayi Yang, Tianchen Zhao, Feng Zhou, Xiaohui Song, Guohao Dai, Shengen Yan, Huazhong Yang, Yu Wang
ICLR 2026 · LLM KV Cache Mixed Precision · Paper · Code · Scholar
Progressively quantizes KV caches with mixed precision for long chain-of-thought inference.
Early diffusion PTQ addresses denoising-step sensitivity; later work extends to diffusion transformers, low-rank outlier handling and video generation.
Post-training Quantization on Diffusion Models
Yuzhang Shang, Zhihang Yuan, Bin Xie, Bingzhe Wu, Yan Yan
CVPR 2023 · Diffusion PTQ · Paper · Code · Scholar
Adapts post-training quantization to diffusion model inference.
Q-diffusion: Quantizing Diffusion Models
Xiuyu Li, Yijiang Liu, Long Lian, Huanrui Yang, Zhen Dong, Daniel Kang, Shanghang Zhang, Kurt Keutzer
ICCV 2023 · Diffusion PTQ · Paper · Code · Scholar
Quantizes diffusion models to reduce the cost of iterative generation.
PTQD: Accurate Post-Training Quantization for Diffusion Models
Yefei He, Luping Liu, Jing Liu, Weijia Wu, Hong Zhou, Bohan Zhuang
NeurIPS 2023 · Diffusion PTQ Error Handling · Paper · Code · Scholar
Targets accurate diffusion generation through post-training quantization error handling.
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
Tianchen Zhao, Tongcheng Fang, Haofeng Huang, Rui Wan, Widyadewi Soedarmadji, Enshu Liu, Shiyao Li, Zinan Lin, Guohao Dai, Shengen Yan, Huazhong Yang, Xuefei Ning, Yu Wang
ICLR 2025 · Diffusion Transformer Image + Video Low-Bit · Paper · Code · Scholar
Quantizes diffusion transformers for both image and video generation.
SVDQuant: Absorbing Outliers by Low-Rank Component for 4-Bit Diffusion Models
Muyang Li, Yujun Lin, Zhekai Zhang, Tianle Cai, Xiuyu Li, Junxian Guo, Enze Xie, Chenlin Meng, Jun-Yan Zhu, Song Han
ICLR 2025 · Diffusion 4-Bit Low-Rank · Paper · Code · Scholar
Absorbs outliers into a low-rank component to support 4-bit diffusion models.
BinaryDM: Accurate Weight Binarization for Efficient Diffusion Models
Xingyu Zheng, Xianglong Liu, Haotong Qin, Xudong Ma, Mingyuan Zhang, Haojie Hao, Jiakai Wang, Zixiang Zhao, Jinyang Guo, Michele Magno
ICLR 2025 · Diffusion Binary Weights · Paper · Code · Scholar
Binarizes diffusion model weights for efficient generation.
Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers
Weilun Feng, Chuanguang Yang, Haotong Qin, Xiangqi Li, Yu Wang, Zhulin An, Libo Huang, Boyu Diao, Zixiang Zhao, Yongjun Xu, Michele Magno
ICML 2025 · Video Diffusion Quantization Distillation · Paper · Code · Scholar
Combines quantization and distillation for video-generation diffusion transformers.
S²Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token Distillation
Weilun Feng, Haotong Qin, Chuanguang Yang, Xiangqi Li, Han Yang, Yuqi Li, Zhulin An, Libo Huang, Michele Magno, Yongjun Xu
NeurIPS 2025 · Video Diffusion Quantization Distillation · Paper · Code · Scholar
Uses salient data and sparse-token distillation to improve quantized video diffusion transformers.
QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention Sparsification
Weilun Feng, Chuanguang Yang, Haotong Qin, Mingqiang Wu, Yuqi Li, Xiangqi Li, Zhulin An, Libo Huang, Yulun Zhang, Michele Magno, Yongjun Xu
ICLR 2026 · Video Diffusion Quantization Attention Sparsity · Paper · Code · Scholar
Combines model quantization and attention sparsification to compress video diffusion transformers.
Vision-language models and selective state space models introduce quantization sensitivities beyond those of language-only transformers.
Q-VLM: Post-training Quantization for Large Vision-Language Models
Changyuan Wang, Ziwei Wang, Xiuwei Xu, Yansong Tang, Jie Zhou, Jiwen Lu
NeurIPS 2024 · VLM PTQ Cross-Layer Dependency · Paper · Code · Scholar
Uses cross-layer dependencies to guide block partitioning and quantization of vision-language models.
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
Hung-Yueh Chiang, Chi-Chih Chang, Natalia Frumkin, Kai-Chiang Wu, Mohamed S. Abdelfattah, Diana Marculescu
ICML 2025 · Mamba State Space Models PTQ W4A8 / W8A8 · Paper · Code · Scholar
Uses channel clustering and state-group quantization to accommodate the sensitivity of Mamba’s selective state-space computations.
Vision methods and deployment systems connect quantizer design to integer kernels, memory movement and hardware costs.
FINN: A Framework for Fast, Scalable Binarized Neural Network Inference
Yaman Umuroglu, Nicholas J. Fraser, Giulio Gambardella, Michaela Blott, Philip Leong, Magnus Jahre, Kees Vissers
FPGA 2017 · Binary Networks FPGA Inference · Paper · Code · Scholar
Provides a framework for fast, scalable binarized neural network inference on FPGA hardware.
HAQ: Hardware-Aware Automated Quantization with Mixed Precision
Kuan Wang, Zhijian Liu, Yujun Lin, Ji Lin, Song Han
CVPR 2019 · CNN Mixed Precision Hardware-Aware · Paper · Code · Scholar
Automates mixed-precision quantization with hardware deployment costs in view.
BiPointNet: Binary Neural Network for Point Clouds
Haotong Qin, Zhongang Cai, Mingyuan Zhang, Yifu Ding, Haiyu Zhao, Shuai Yi, Xianglong Liu, Hao Su
ICLR 2021 · Point Clouds Binary QAT 1-Bit · Paper · Code · Scholar
Uses entropy-maximizing aggregation and layer-wise scale recovery to address feature homogenization and scale distortion in binary point-cloud networks.
PTQ4ViT: Post-Training Quantization for Vision Transformers with Twin Uniform Quantization
Zhihang Yuan, Chenhao Xue, Yiqi Chen, Qiang Wu, Guangyu Sun
ECCV 2022 · Vision Transformer PTQ · Paper · Code · Scholar
Uses twin uniform quantization to support post-training compression of vision transformers.
QuantSR: Accurate Low-bit Quantization for Efficient Image Super-Resolution
Haotong Qin, Yulun Zhang, Yifu Ding, Yifan Liu, Xianglong Liu, Martin Danelljan, Fisher Yu
NeurIPS 2023 · Super-Resolution QAT 2–4 Bit · Paper · Code · Scholar
Combines a redistribution-driven learnable quantizer with a depth-dynamic architecture for accurate low-bit image super-resolution.
LUT-GEMM: Quantized Matrix Multiplication based on LUTs for Efficient Inference in Large-Scale Generative Language Models
Gunho Park, Baeseong Park, Minsub Kim, Sungjae Lee, Jeonghoon Kim, Beomseok Kwon, Se Jung Kwon, Byeongwook Kim, Youngjoo Lee, Dongsoo Lee
ICLR 2024 · LLM Quantized Matrix Multiplication Lookup Tables · Paper · Scholar
Uses lookup tables for efficient quantized matrix multiplication in generative language models.
Quant-LLM: Accelerating the Serving of Large Language Models via FP6-Centric Algorithm-System Co-Design on Modern GPUs
Haojun Xia, Zhen Zheng, Xiaoxia Wu, Shiyang Chen, Zhewei Yao, Stephen Youn, Arash Bakhtiari, Michael Wyatt, Donglin Zhuang, Zhongzhu Zhou, Olatunji Ruwase, Yuxiong He, Shuaiwen Leon Song
USENIX ATC 2024 · LLM FP6 GPU Kernels · Paper · Code · Scholar
Uses TC-FPx kernels to support non-power-of-two weight formats efficiently on GPUs; the codebase is also known as FP6-LLM.
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, Song Han
MLSys 2025 · LLM W4A8KV4 GPU Serving · Paper · Code · Scholar
Co-designs progressive quantization, attention and GPU kernels to turn reduced precision into serving throughput.
Low-bit floating-point and shared-scale formats complement integer quantization. Format design and model calibration are separate choices.
FP8 Formats for Deep Learning
Paulius Micikevicius, Dusan Stosic, Neil Burgess, Marius Cornea, Pradeep Dubey, Richard Grisenthwaite, Sangwon Ha, Alexander Heinecke, Patrick Judd, John Kamalu, Naveen Mellempudi, Stuart Oberman, Mohammad Shoeybi, Michael Siu, Hao Wu
arXiv 2022 · FP8 E4M3 / E5M2 Training + Inference · Paper · Scholar
Defines complementary FP8 encodings and evaluates their use in neural network training and inference.
Microscaling Data Formats for Deep Learning
Bita Darvish Rouhani, Ritchie Zhao, Ankit More, Mathew Hall, Alireza Khodamoradi, Summer Deng, Dhruv Choudhary, Marius Cornea, Eric Dellinger, Kristof Denolf, Stosic Dusan, Venmugil Elango, Maximilian Golub, Alexander Heinecke, Phil James-Roxby, Dharmesh Jani, Gaurav Kolhe, Martin Langhammer, Ada Li, Levi Melnick, Maral Mesmakhosroshahi, Andres Rodriguez, Michael Schulte, Rasoul Shafipour, Lei Shao, Michael Siu, Pradeep Dubey, Paulius Micikevicius, Maxim Naumov, Colin Verrilli, Ralph Wittig, Doug Burger, Eric Chung
arXiv 2023 · MX Formats Block Scaling Training + Inference · Paper · Code · Scholar
Combines shared block scales with narrow element formats to balance numerical range and hardware efficiency.
LLM-FP4: 4-Bit Floating-Point Quantized Transformers
Shih-yang Liu, Zechun Liu, Xijie Huang, Pingcheng Dong, Kwang-Ting Cheng
EMNLP 2023 · LLM PTQ FP4 · Paper · Code · Scholar
Searches exponent configurations and quantization parameters to handle weight and activation range differences in 4-bit floating point.
Quantization can reduce saved activations, optimizer states, gradient communication or training arithmetic; each targets a different part of the training cost.
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, Milan Vojnovic
NeurIPS 2017 · Training Gradients Communication · Paper · Scholar
Uses randomized gradient quantization with convergence guarantees to trade communication bandwidth against estimator variance.
ActNN: Reducing Training Memory Footprint via 2-Bit Activation Compressed Training
Jianfei Chen, Lianmin Zheng, Zhewei Yao, Dequan Wang, Ion Stoica, Michael Mahoney, Joseph Gonzalez
ICML 2021 · Training Activations 2-Bit · Paper · Code · Scholar
Compresses saved activations to reduce the memory footprint of neural network training.
8-bit Optimizers via Block-wise Quantization
Tim Dettmers, Mike Lewis, Sam Shleifer, Luke Zettlemoyer
ICLR 2022 · Training Optimizer States 8-Bit · Paper · Code · Scholar
Uses block-wise quantization to reduce optimizer-state memory.
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
Jinda Jia, Cong Xie, Hanlin Lu, Daoce Wang, Hao Feng, Chengming Zhang, Baixi Sun, Haibin Lin, Zhi Zhang, Xin Liu, Dingwen Tao
NeurIPS 2024 · LLM Training Communication 4-Bit · Paper · Code · Scholar
Targets 4-bit communication quantization in sharded data-parallel LLM training.
Optimizing Large Language Model Training Using FP4 Quantization
Ruizhe Wang, Yeyun Gong, Xiao Liu, Guoshuai Zhao, Ziyue Yang, Baining Guo, Zhengjun Zha, Peng Cheng
ICML 2025 · LLM Training FP4 Gradient Estimation · Paper · Scholar
Combines differentiable quantization estimation with outlier handling to stabilize FP4 LLM training.
Choose by evaluation scope: deployment reproducibility, binary networks, LLM capabilities or robustness. Expand a resource below for authors, figures and citation details.
| Resource | What it covers |
|---|---|
| MQBench: Towards Reproducible and Deployable Model Quantization Benchmark NeurIPS 2021 Datasets and Benchmarks Code · Scholar | QAT + deployment Compares quantization algorithms under reproducible settings and hardware backend constraints. |
| BiBench: Benchmarking and Analyzing Network Binarization ICML 2023 Code · Scholar | Binary networks Compares binarization methods across tasks, architectures and deployment settings. |
| Evaluating Quantized Large Language Models ICML 2024 Code · Scholar | Weights, activations + KV cache Evaluates 11 model families on basic NLP, emergent abilities, trustworthiness, dialogue and long-context tasks. |
| LLMC: Benchmarking Large Language Model Quantization with a Versatile Compression Toolkit EMNLP 2024 Industry Track Code · Scholar | LLM toolkit Compares calibration data, method pipelines and quantization configurations; the toolkit is now LightCompress. |
| An empirical study of LLaMA3 quantization: from LLMs to MLLMs Visual Intelligence 2024 Code · Scholar | LLMs + multimodal Examines low-bit behavior across LLaMA3 language and multimodal models. |
| An Empirical Study of Qwen3 Quantization Visual Intelligence 2026 Code · Scholar | Dense + MoE LLMs Studies quantization across Qwen3 model sizes, architectures and reasoning settings. |
| RobustMQ: Benchmarking Robustness of Quantized Models Visual Intelligence 2023 Scholar | Model robustness Tests quantized models beyond clean accuracy, including robustness under input perturbations. |
Yuhang Li, Mingzhu Shen, Jian Ma, Yan Ren, Mingxin Zhao, Qi Zhang, Ruihao Gong, Fengwei Yu, Junjie Yan
@inproceedings{li2021mqbench,
title={MQBench: Towards Reproducible and Deployable Model Quantization Benchmark},
author={Li, Yuhang and Shen, Mingzhu and Ma, Jian and Ren, Yan and Zhao, Mingxin and Zhang, Qi and Gong, Ruihao and Yu, Fengwei and Yan, Junjie},
booktitle={NeurIPS Datasets and Benchmarks},
year={2021}
}
Haotong Qin, Mingyuan Zhang, Yifu Ding, Aoyu Li, Zhongang Cai, Ziwei Liu, Fisher Yu, Xianglong Liu

@inproceedings{qin2023bibench,
title={BiBench: Benchmarking and Analyzing Network Binarization},
author={Qin, Haotong and Zhang, Mingyuan and Ding, Yifu and Li, Aoyu and Cai, Zhongang and Liu, Ziwei and Yu, Fisher and Liu, Xianglong},
booktitle={International Conference on Machine Learning (ICML)},
year={2023}
}
Shiyao Li, Xuefei Ning, Luning Wang, Tengxuan Liu, Xiangsheng Shi, Shengen Yan, Guohao Dai, Huazhong Yang, Yu Wang
@inproceedings{li2024evaluating,
title={Evaluating Quantized Large Language Models},
author={Li, Shiyao and Ning, Xuefei and Wang, Luning and Liu, Tengxuan and Shi, Xiangsheng and Yan, Shengen and Dai, Guohao and Yang, Huazhong and Wang, Yu},
booktitle={International Conference on Machine Learning},
year={2024},
url={https://proceedings.mlr.press/v235/li24bb.html}
}
Ruihao Gong, Yang Yong, Shiqiao Gu, Yushi Huang, Chengtao Lv, Yunchen Zhang, Dacheng Tao, Xianglong Liu

@inproceedings{gong2024llmc,
title={Llmc: Benchmarking large language model quantization with a versatile compression toolkit},
author={Gong, Ruihao and Yong, Yang and Gu, Shiqiao and Huang, Yushi and Lv, Chengtao and Zhang, Yunchen and Tao, Dacheng and Liu, Xianglong},
booktitle={Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track},
pages={132--152},
year={2024}
}
Wei Huang, Xingyu Zheng, Xudong Ma, Haotong Qin, Chengtao Lv, Hong Chen, Jie Luo, Xiaojuan Qi, Xianglong Liu, Michele Magno

@article{huang2024empirical,
title={An empirical study of llama3 quantization: From llms to mllms},
author={Huang, Wei and Zheng, Xingyu and Ma, Xudong and Qin, Haotong and Lv, Chengtao and Chen, Hong and Luo, Jie and Qi, Xiaojuan and Liu, Xianglong and Magno, Michele},
journal={Visual Intelligence},
volume={2},
number={1},
pages={36},
year={2024},
publisher={Springer}
}
Xingyu Zheng, Yuye Li, Haoran Chu, Yue Feng, Xudong Ma, Zining Wang, Jie Luo, Jinyang Guo, Haotong Qin, Michele Magno, Xianglong Liu

@article{zheng2026empirical,
title={An empirical study of Qwen3 quantization},
author={Zheng, Xingyu and Li, Yuye and Chu, Haoran and Feng, Yue and Ma, Xudong and Wang, Zining and Luo, Jie and Guo, Jinyang and Qin, Haotong and Magno, Michele and Liu, Xianglong},
journal={Visual Intelligence},
volume={4},
pages={11},
year={2026},
doi={10.1007/s44267-026-00114-4}
}
Yisong Xiao, Aishan Liu, Tianyuan Zhang, Haotong Qin, Jinyang Guo, Xianglong Liu

@article{xiao2023robustmq,
title={Robustmq: benchmarking robustness of quantized models},
author={Xiao, Yisong and Liu, Aishan and Zhang, Tianyuan and Qin, Haotong and Guo, Jinyang and Liu, Xianglong},
journal={Visual Intelligence},
volume={1},
number={1},
pages={30},
year={2023},
publisher={Springer}
}
Start with the white paper for practical PTQ/QAT, then choose a survey for broader context or a specific model family. Figures and citation details are available below.
| Resource | What it covers |
|---|---|
| A White Paper on Neural Network Quantization arXiv 2021 Scholar | Practical PTQ + QAT Explains quantizer design, common failure modes and practical post-training and quantization-aware training workflows. |
| A Survey of Quantization Methods for Efficient Neural Network Inference arXiv 2021 Scholar | Foundations + taxonomy Reviews quantization design choices, mixed precision and the trade-offs between model accuracy and efficient inference. |
| Binary Neural Networks: A Survey Pattern Recognition 2020 Scholar | Binary networks Surveys binary network representations, training methods and applications. |
| A Survey of Low-bit Large Language Models: Basics, Systems, and Algorithms Neural Networks 2025 Scholar | LLM algorithms + systems Connects low-bit LLM algorithms with numerical formats and inference systems. |
| Low-bit Model Quantization for Deep Neural Networks: A Survey arXiv 2025 Scholar | Broad low-bit methods Maps low-bit quantization methods across neural network architectures and applications. |
Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, Tijmen Blankevoort
@article{nagel2021white,
title={A White Paper on Neural Network Quantization},
author={Nagel, Markus and Fournarakis, Marios and Amjad, Rana Ali and Bondarenko, Yelysei and van Baalen, Mart and Blankevoort, Tijmen},
journal={arXiv preprint arXiv:2106.08295},
year={2021}
}
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, Kurt Keutzer
@article{gholami2021survey,
title={A Survey of Quantization Methods for Efficient Neural Network Inference},
author={Gholami, Amir and Kim, Sehoon and Dong, Zhen and Yao, Zhewei and Mahoney, Michael W. and Keutzer, Kurt},
journal={arXiv preprint arXiv:2103.13630},
year={2021}
}
Haotong Qin, Ruihao Gong, Xianglong Liu, Xiao Bai, Jingkuan Song, Nicu Sebe

@article{Qin:pr20_bnn_survey,
title = "Binary neural networks: A survey",
author = "Haotong Qin and Ruihao Gong and Xianglong Liu and Xiao Bai and Jingkuan Song and Nicu Sebe",
journal = "Pattern Recognition",
volume = "105",
pages = "107281",
year = "2020"
}
Ruihao Gong, Yifu Ding, Zining Wang, Chengtao Lv, Xingyu Zheng, Jinyang Du, Yang Yong, Shiqiao Gu, Haotong Qin, Jinyang Guo, Dahua Lin, Michele Magno, Xianglong Liu

@article{gong2025survey,
title={A survey of low-bit large language models: Basics, systems, and algorithms},
author={Gong, Ruihao and Ding, Yifu and Wang, Zining and Lv, Chengtao and Zheng, Xingyu and Du, Jinyang and Yong, Yang and Gu, Shiqiao and Qin, Haotong and Guo, Jinyang and Lin, Dahua and Magno, Michele and Liu, Xianglong},
journal={Neural Networks},
pages={107856},
year={2025}
}
Kai Liu, Qian Zheng, Kaiwen Tao, Zhiteng Li, Haotong Qin, Wenbo Li, Yong Guo, Xianglong Liu, Linghe Kong, Guihai Chen, Yulun Zhang, Xiaokang Yang

@article{liu2025low,
title={Low-bit Model Quantization for Deep Neural Networks: A Survey},
author={Liu, Kai and Zheng, Qian and Tao, Kaiwen and Li, Zhiteng and Qin, Haotong and Li, Wenbo and Guo, Yong and Liu, Xianglong and Kong, Linghe and Chen, Guihai and Zhang, Yulun and Yang, Xiaokang},
journal={arXiv preprint arXiv:2505.05530},
year={2025}
}
All paper titles and links are kept in this README. Published work is grouped by venue year where verified; otherwise the recorded preprint year is used. Representative works, benchmarks and surveys also appear here for chronological browsing. Within each year, entries are grouped by conference or journal, with preprints at the end.
Truncated — view the full README on GitHub.
(top 30 of 34)