[ACL-2026 Findings] Implementation for HiPrune, a training-free visual token pruning method for VLM acceleration.
Python
61
17 commits
updated Apr 29, 2026
Vision-Language Models (VLMs) encode images into lengthy sequences of visual tokens, leading to excessive computational overhead and limited inference efficiency. While prior efforts prune or merge tokens to address this issue, they often rely on special tokens (e.g., CLS) or require task-specific training, hindering scalability across architectures. In this paper, we propose HiPrune, a training-free and model-agnostic token Pruning framework that exploits the Hierarchical attention structure within vision encoders. We identify that middle layers attend to object-centric regions, while deep layers capture global contextual features. Based on this observation, HiPrune selects three types of informative tokens: (1) Anchor tokens with high attention in object-centric layers, (2) Buffer tokens adjacent to anchors for spatial continuity, and (3) Register tokens with strong attention in deep layers for global summarization. Our method requires no retraining and integrates seamlessly with any ViT-based VLM. Extensive experiments on LLaVA-1.5, LLaVA-NeXT, and Qwen2.5-VL demonstrate that HiPrune achieves state-of-the-art pruning performance, preserving up to 99.3% task accuracy with only 33.3% tokens, and maintaining 99.5% accuracy with just 11.1% tokens. Meanwhile, it reduces inference FLOPs and latency by up to 9$\times$, showcasing strong generalization across models and tasks.



To better clarify our method, we first define three types of tokens:
HiPrune retains these three types of tokens and discards the rest before the LLM. Token identification is also an ordered process, from anchor tokens to buffer tokens and finally register tokens, guided by the hierarchical attention from the vision encoder.
Environment Setup. Since LLaVA and Qwen require different versions of transformers, so please setup the corresponding environment before running them.
bash setup_llava.sh
bash setup_qwen.sh
Data Access. Before starting your evaluation, please login with your huggingface token to get access to some datasets.
huggingface-cli login
Before starting evaluation on LLaVA-1.5, please follow its official instruction to prepare data for MMB, MMBCN, textVQA, and VQAv2.
bash bench_llava.sh
bash bench_llava_next.sh
python bench_sys.py
bash flops.py
bash bench_qwen.sh
If you want to change hyperparameters, you can simply set environment variables in this table.
| Model | Environment Variables | Range | Default | Denote | Description |
|---|---|---|---|---|---|
| LLaVA-1.5 | HIPRUNE_RETENTION | 1-576 | 192/128/64 | $N'$ | Token budget |
| HIPRUNE_ALPHA | 0-1 | 0.1 | $\alpha$ | Proportation of anchor and buffer tokens | |
| HIPRUNE_OBJECT_LAYER | 1-24 | 9 | $l$ | Object layer to choose anchor and buffer tokens | |
| LLaVA-NeXT | HIPRUNE_RETENTION | 1-2880 | 640/320/160 | $N'$ | Token budget |
| HIPRUNE_ALPHA | 0-1 | 0.1 | $\alpha$ | Proportation of anchor and buffer tokens | |
| HIPRUNE_OBJECT_LAYER | 1-24 | 9 | $l$ | Object layer to choose anchor and buffer tokens | |
| Qwen2.5-VL | HIPRUNE_QWEN_RETENTION | 0-1 | 0.334/0.223/0.112 | $\frac{N'}{N}$ | Token retention ratio |
| HIPRUNE_ALPHA | 0-1 | 0.1 | $\alpha$ | Proportation of anchor and buffer tokens | |
| HIPRUNE_OBJECT_LAYER | 1-24 | 16 | $l$ | Object layer to choose anchor and buffer tokens |
This repository is built on LLaVA, FasterVLM, lmms-eval. Acknowledge their outstanding work!
If you found our work helpful, please consider leaving a star ⭐ and citing our work.
@inproceedings{liu2026hiprunesapp,
title={HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models (Student Abstract)},
author={Liu, Jizhihui and Zhu, Guangdao and Du, Feiyi},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
volume={40},
number={48},
pages={41275--41277},
year={2026}
}
@inproceedings{
liu2026hiprune,
title={HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models},
author={Jizhihui Liu, Guangdao Zhu, Feiyi Du, Niu Lian, Jun Li, Bin Chen, Weili Guan, Yaowei Wang},
booktitle={The 64th Annual Meeting of the Association for Computational Linguistics},
year={2026}
}
[ACL-2026 Findings] Implementation for HiPrune, a training-free visual token pruning method for VLM acceleration.
Python
61
17 commits
updated Apr 29, 2026
Vision-Language Models (VLMs) encode images into lengthy sequences of visual tokens, leading to excessive computational overhead and limited inference efficiency. While prior efforts prune or merge tokens to address this issue, they often rely on special tokens (e.g., CLS) or require task-specific training, hindering scalability across architectures. In this paper, we propose HiPrune, a training-free and model-agnostic token Pruning framework that exploits the Hierarchical attention structure within vision encoders. We identify that middle layers attend to object-centric regions, while deep layers capture global contextual features. Based on this observation, HiPrune selects three types of informative tokens: (1) Anchor tokens with high attention in object-centric layers, (2) Buffer tokens adjacent to anchors for spatial continuity, and (3) Register tokens with strong attention in deep layers for global summarization. Our method requires no retraining and integrates seamlessly with any ViT-based VLM. Extensive experiments on LLaVA-1.5, LLaVA-NeXT, and Qwen2.5-VL demonstrate that HiPrune achieves state-of-the-art pruning performance, preserving up to 99.3% task accuracy with only 33.3% tokens, and maintaining 99.5% accuracy with just 11.1% tokens. Meanwhile, it reduces inference FLOPs and latency by up to 9$\times$, showcasing strong generalization across models and tasks.



To better clarify our method, we first define three types of tokens:
HiPrune retains these three types of tokens and discards the rest before the LLM. Token identification is also an ordered process, from anchor tokens to buffer tokens and finally register tokens, guided by the hierarchical attention from the vision encoder.
Environment Setup. Since LLaVA and Qwen require different versions of transformers, so please setup the corresponding environment before running them.
bash setup_llava.sh
bash setup_qwen.sh
Data Access. Before starting your evaluation, please login with your huggingface token to get access to some datasets.
huggingface-cli login
Before starting evaluation on LLaVA-1.5, please follow its official instruction to prepare data for MMB, MMBCN, textVQA, and VQAv2.
bash bench_llava.sh
bash bench_llava_next.sh
python bench_sys.py
bash flops.py
bash bench_qwen.sh
If you want to change hyperparameters, you can simply set environment variables in this table.
| Model | Environment Variables | Range | Default | Denote | Description |
|---|---|---|---|---|---|
| LLaVA-1.5 | HIPRUNE_RETENTION | 1-576 | 192/128/64 | $N'$ | Token budget |
| HIPRUNE_ALPHA | 0-1 | 0.1 | $\alpha$ | Proportation of anchor and buffer tokens | |
| HIPRUNE_OBJECT_LAYER | 1-24 | 9 | $l$ | Object layer to choose anchor and buffer tokens | |
| LLaVA-NeXT | HIPRUNE_RETENTION | 1-2880 | 640/320/160 | $N'$ | Token budget |
| HIPRUNE_ALPHA | 0-1 | 0.1 | $\alpha$ | Proportation of anchor and buffer tokens | |
| HIPRUNE_OBJECT_LAYER | 1-24 | 9 | $l$ | Object layer to choose anchor and buffer tokens | |
| Qwen2.5-VL | HIPRUNE_QWEN_RETENTION | 0-1 | 0.334/0.223/0.112 | $\frac{N'}{N}$ | Token retention ratio |
| HIPRUNE_ALPHA | 0-1 | 0.1 | $\alpha$ | Proportation of anchor and buffer tokens | |
| HIPRUNE_OBJECT_LAYER | 1-24 | 16 | $l$ | Object layer to choose anchor and buffer tokens |
This repository is built on LLaVA, FasterVLM, lmms-eval. Acknowledge their outstanding work!
If you found our work helpful, please consider leaving a star ⭐ and citing our work.
@inproceedings{liu2026hiprunesapp,
title={HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models (Student Abstract)},
author={Liu, Jizhihui and Zhu, Guangdao and Du, Feiyi},
booktitle={Proceedings of the AAAI Conference on Artificial Intelligence},
volume={40},
number={48},
pages={41275--41277},
year={2026}
}
@inproceedings{
liu2026hiprune,
title={HiPrune: Hierarchical Attention for Efficient Token Pruning in Vision-Language Models},
author={Jizhihui Liu, Guangdao Zhu, Feiyi Du, Niu Lian, Jun Li, Bin Chen, Weili Guan, Yaowei Wang},
booktitle={The 64th Annual Meeting of the Association for Computational Linguistics},
year={2026}
}