[ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
See the codePairwise Rotation Quantization for Efficient Reasoning LLM Inference
State-of-the-art INT4 quantization for LLMs. ParoQuant uses learned pairwise rotations to suppress weight outliers, closing the accuracy gap with FP16 while running at near-AWQ speed. Supports NVIDIA GPUs (vLLM, Transformers) and Apple Silicon (MLX).
On Apple Silicon, we recommend serving ParoQuant models using oMLX. Refer to the docs for more details.
# NVIDIA GPU (CUDA 12.9)
pip install "paroquant[vllm]"
# NVIDIA GPU (CUDA 13.0)
pip install "paroquant[vllm]" "vllm==0.19.1" \
--extra-index-url https://wheels.vllm.ai/0.19.1/cu130 \
--extra-index-url https://download.pytorch.org/whl/cu130
# Apple Silicon
pip install "paroquant[mlx]"
Pick a model from our Hugging Face collection:
export MODEL=z-lab/Qwen3.5-4B-PARO
python -m paroquant.cli.chat --model $MODEL
For vLLM, you can directly use vllm serve to serve ParoQuant models:
vllm serve $MODEL --port 8000
For other frameworks:
python -m paroquant.cli.serve --model $MODEL --port 8000
For MLX, add --vlm if you wish to load the VLM components and use the model's multimodal features. For vLLM, VLM components are loaded by default and can be skipped with the server argument --language-model-only.
[!NOTE] The following commands map the local cache directory to the container in order to persist kernel cache across runs. Remove
-v ...to disable this behaviour.
# Interactive chat
docker run --pull=always --rm -it --gpus all --ipc=host \
-v $HOME/.cache/paroquant:/root/.cache/paroquant \
ghcr.io/z-lab/paroquant:chat --model $MODEL
# API server (port 8000)
docker run --pull=always --rm -it --gpus all --ipc=host -p 8000:8000 \
-v $HOME/.cache/paroquant:/root/.cache/paroquant \
ghcr.io/z-lab/paroquant:serve --model $MODEL
All models are available on Hugging Face. Swap the model name in the commands above to try any of them.
Qwen3.8
| Model | Checkpoint |
|---|---|
| Qwen3.8-27B | z-lab/Qwen3.8-27B-PARO |
Gemma 4
| Model | Checkpoint |
|---|---|
| gemma-4-31B-it | z-lab/gemma-4-31B-it-PARO |
| gemma-4-12B-it | z-lab/gemma-4-12B-it-PARO |
| gemma-4-26B-A4B-it | z-lab/gemma-4-26B-A4B-it-PARO |
| gemma-4-E4B-it | z-lab/gemma-4-E4B-it-PARO |
| gemma-4-E2B-it | z-lab/gemma-4-E2B-it-PARO |
Qwen3.6
| Model | Checkpoint |
|---|---|
| Qwen3.6-27B | z-lab/Qwen3.6-27B-PARO |
| Qwen3.6-35B-A3B | z-lab/Qwen3.6-35B-A3B-PARO |
Qwen3.5
| Model | Checkpoint |
|---|---|
| Qwen3.5-0.8B | z-lab/Qwen3.5-0.8B-PARO |
| Qwen3.5-2B | z-lab/Qwen3.5-2B-PARO |
| Qwen3.5-4B | z-lab/Qwen3.5-4B-PARO |
| Qwen3.5-9B | z-lab/Qwen3.5-9B-PARO |
| Qwen3.5-27B | z-lab/Qwen3.5-27B-PARO |
| Qwen3.5-35B-A3B | z-lab/Qwen3.5-35B-A3B-PARO |
Qwen3
| Model | Checkpoint |
|---|---|
| Qwen3-0.6B | z-lab/Qwen3-0.6B-PARO |
| Qwen3-1.7B | z-lab/Qwen3-1.7B-PARO |
| Qwen3-4B | z-lab/Qwen3-4B-PARO |
| Qwen3-8B | z-lab/Qwen3-8B-PARO |
| Qwen3-14B | z-lab/Qwen3-14B-PARO |
Llama
| Model | Checkpoint |
|---|---|
| Llama-2-7B | z-lab/Llama-2-7b-hf-PARO |
| Llama-3-8B | z-lab/Meta-Llama-3-8B-PARO |
| Llama-3.1-8B-Instruct | z-lab/Llama-3.1-8B-Instruct-PARO |
Want a model that's not listed? Open an issue and let us know.
[!NOTE] The main branch of this repository is under active development, and reproducibility is not guaranteed. Please use the
legacybranch to reproduce results from the paper.
git clone https://github.com/z-lab/paroquant && cd paroquant
pip install -e ".[optim,eval]"
# 1. Optimize rotation parameters
experiments/optimize/4bit.sh Qwen/Qwen3-8B
# 2. Export to HF checkpoint (--mode real for INT4, --mode pseudo for FP16)
python -m paroquant.cli.convert \
--model Qwen/Qwen3-8B \
--result-dir output/Qwen3-8B \
--output-path models/Qwen3-8B-PARO
| Image | Purpose |
|---|---|
ghcr.io/z-lab/paroquant:chat | Interactive chat |
ghcr.io/z-lab/paroquant:chat-cu129 | Interactive chat (CUDA 12.9) |
ghcr.io/z-lab/paroquant:serve | OpenAI-compatible API server |
ghcr.io/z-lab/paroquant:latest | Optimization & evaluation |
ghcr.io/z-lab/paroquant:eval | Reasoning task evaluation |
@inproceedings{liang2026paroquant,
title = {{ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference}},
author = {Liang, Yesheng and Chen, Haisheng and Zhang, Zihan and Han, Song and Liu, Zhijian},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}
530 followers · starred Mar 2026
404 followers · starred Mar 2026
216 followers · starred May 2026
25 followers · starred Apr 2026
Python
85.7%
Shell
5.5%
Cuda
5.3%
Jinja
1.4%
Dockerfile
1.1%
[ICLR 2026] ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
See the codePairwise Rotation Quantization for Efficient Reasoning LLM Inference
State-of-the-art INT4 quantization for LLMs. ParoQuant uses learned pairwise rotations to suppress weight outliers, closing the accuracy gap with FP16 while running at near-AWQ speed. Supports NVIDIA GPUs (vLLM, Transformers) and Apple Silicon (MLX).
On Apple Silicon, we recommend serving ParoQuant models using oMLX. Refer to the docs for more details.
# NVIDIA GPU (CUDA 12.9)
pip install "paroquant[vllm]"
# NVIDIA GPU (CUDA 13.0)
pip install "paroquant[vllm]" "vllm==0.19.1" \
--extra-index-url https://wheels.vllm.ai/0.19.1/cu130 \
--extra-index-url https://download.pytorch.org/whl/cu130
# Apple Silicon
pip install "paroquant[mlx]"
Pick a model from our Hugging Face collection:
export MODEL=z-lab/Qwen3.5-4B-PARO
python -m paroquant.cli.chat --model $MODEL
For vLLM, you can directly use vllm serve to serve ParoQuant models:
vllm serve $MODEL --port 8000
For other frameworks:
python -m paroquant.cli.serve --model $MODEL --port 8000
For MLX, add --vlm if you wish to load the VLM components and use the model's multimodal features. For vLLM, VLM components are loaded by default and can be skipped with the server argument --language-model-only.
[!NOTE] The following commands map the local cache directory to the container in order to persist kernel cache across runs. Remove
-v ...to disable this behaviour.
# Interactive chat
docker run --pull=always --rm -it --gpus all --ipc=host \
-v $HOME/.cache/paroquant:/root/.cache/paroquant \
ghcr.io/z-lab/paroquant:chat --model $MODEL
# API server (port 8000)
docker run --pull=always --rm -it --gpus all --ipc=host -p 8000:8000 \
-v $HOME/.cache/paroquant:/root/.cache/paroquant \
ghcr.io/z-lab/paroquant:serve --model $MODEL
All models are available on Hugging Face. Swap the model name in the commands above to try any of them.
Qwen3.8
| Model | Checkpoint |
|---|---|
| Qwen3.8-27B | z-lab/Qwen3.8-27B-PARO |
Gemma 4
| Model | Checkpoint |
|---|---|
| gemma-4-31B-it | z-lab/gemma-4-31B-it-PARO |
| gemma-4-12B-it | z-lab/gemma-4-12B-it-PARO |
| gemma-4-26B-A4B-it | z-lab/gemma-4-26B-A4B-it-PARO |
| gemma-4-E4B-it | z-lab/gemma-4-E4B-it-PARO |
| gemma-4-E2B-it | z-lab/gemma-4-E2B-it-PARO |
Qwen3.6
| Model | Checkpoint |
|---|---|
| Qwen3.6-27B | z-lab/Qwen3.6-27B-PARO |
| Qwen3.6-35B-A3B | z-lab/Qwen3.6-35B-A3B-PARO |
Qwen3.5
| Model | Checkpoint |
|---|---|
| Qwen3.5-0.8B | z-lab/Qwen3.5-0.8B-PARO |
| Qwen3.5-2B | z-lab/Qwen3.5-2B-PARO |
| Qwen3.5-4B | z-lab/Qwen3.5-4B-PARO |
| Qwen3.5-9B | z-lab/Qwen3.5-9B-PARO |
| Qwen3.5-27B | z-lab/Qwen3.5-27B-PARO |
| Qwen3.5-35B-A3B | z-lab/Qwen3.5-35B-A3B-PARO |
Qwen3
| Model | Checkpoint |
|---|---|
| Qwen3-0.6B | z-lab/Qwen3-0.6B-PARO |
| Qwen3-1.7B | z-lab/Qwen3-1.7B-PARO |
| Qwen3-4B | z-lab/Qwen3-4B-PARO |
| Qwen3-8B | z-lab/Qwen3-8B-PARO |
| Qwen3-14B | z-lab/Qwen3-14B-PARO |
Llama
| Model | Checkpoint |
|---|---|
| Llama-2-7B | z-lab/Llama-2-7b-hf-PARO |
| Llama-3-8B | z-lab/Meta-Llama-3-8B-PARO |
| Llama-3.1-8B-Instruct | z-lab/Llama-3.1-8B-Instruct-PARO |
Want a model that's not listed? Open an issue and let us know.
[!NOTE] The main branch of this repository is under active development, and reproducibility is not guaranteed. Please use the
legacybranch to reproduce results from the paper.
git clone https://github.com/z-lab/paroquant && cd paroquant
pip install -e ".[optim,eval]"
# 1. Optimize rotation parameters
experiments/optimize/4bit.sh Qwen/Qwen3-8B
# 2. Export to HF checkpoint (--mode real for INT4, --mode pseudo for FP16)
python -m paroquant.cli.convert \
--model Qwen/Qwen3-8B \
--result-dir output/Qwen3-8B \
--output-path models/Qwen3-8B-PARO
| Image | Purpose |
|---|---|
ghcr.io/z-lab/paroquant:chat | Interactive chat |
ghcr.io/z-lab/paroquant:chat-cu129 | Interactive chat (CUDA 12.9) |
ghcr.io/z-lab/paroquant:serve | OpenAI-compatible API server |
ghcr.io/z-lab/paroquant:latest | Optimization & evaluation |
ghcr.io/z-lab/paroquant:eval | Reasoning task evaluation |
@inproceedings{liang2026paroquant,
title = {{ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference}},
author = {Liang, Yesheng and Chen, Haisheng and Zhang, Zihan and Han, Song and Liu, Zhijian},
booktitle = {International Conference on Learning Representations (ICLR)},
year = {2026}
}
530 followers · starred Mar 2026
404 followers · starred Mar 2026
216 followers · starred May 2026
25 followers · starred Apr 2026
Python
85.7%
Shell
5.5%
Cuda
5.3%
Jinja
1.4%
Dockerfile
1.1%