[ACL2026 Main] Data & Code of "Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods"
Python
37
37 commits
updated Apr 9, 2026
Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
Chenfei Liao1,2,6 Wensong Wang3,2 Zichen Wen2,5 Xu Zheng1,4,6 Yiyu Wang2 Haocong He2
Yuanhuiyi Lyu1,6 Lutao Jiang1,6 Xin Zou1,6 Yuqian Fu4 Bin Ren7,8,4 Linfeng Zhang2,π§ Xuming Hu1,6,π§
1Hong Kong University of Science and Technology (Guangzhou) 2Shanghai Jiao Tong University
3Northeastern University 4INSAIT, Sofia University βSt. Kliment Ohridskiβ
5Shanghai AI Laboratory 6Hong Kong University of Science and Technology
7University of Pisa 8University of Trento
Recent efforts to accelerate inference in Multimodal Large Language Models (MLLMs) have largely focused on visual token compression. The effectiveness of these methods is commonly evaluated by measuring the accuracy drop on existing MLLM benchmarks before and after compression. However, these benchmarks are originally designed to assess general perception and reasoning abilities, rather than the specific challenges posed by visual token compression, leading to a fundamental task mismatch.
In this work, we uncover a counterintuitive yet consistent phenomenon: simple image downsampling outperforms many advanced visual token compression methods across multiple widely used benchmarks.
Through a comprehensive empirical study spanning eight popular benchmarks and multiple state-of-the-art compression techniques, we show that (i) current benchmarks contain substantial noise (task-irrelevant samples) for evaluating visual token compression, and (ii) downsampling can act as an effective data filter that distinguishes between simple and difficult samples with respect to compression sensitivity.
Motivated by these findings, we propose VTC-Bench, an evaluation framework that explicitly leverages downsampling as a discriminator to denoise existing benchmarks, enabling a fairer and more meaningful additional assessment of visual token compression methods.
Some recent MLLMs, such as Qwen2-VL and Qwen2.5-VL, natively support inputs of varying resolutions. A trivial yet efficient method to handle high-resolution images is to simply downsample them to a lower resolution. However, most token compression methods for MLLMs choose to adaptively drop useless tokens or merge similar tokens instead of directly downsampling the original image, which theoretically should be more intelligent.
Surprisingly, we find that image downsampling consistently exceeds other sophisticated methods under some settings. Based on comprehensive experiments, we propose a bold hypothesis:
Some data in the existing benchmarks is overly simplistic and irrelevant to evaluating visual token compression methods, leading to the unreasonable phenomenon that even the downsampling method is sufficient to deal with the visual token compression task.
To validate this, we design a data-centric analysis using downsampling as a discriminator. We identify two crucial findings:
Based on these findings, we propose VTC-Bench, a new evaluation framework specifically designed to optimize and denoise current existing benchmarks. By explicitly distinguishing between βsimpleβ and βdifficultβ samples through downsampling, VTC-Bench adaptively selects "difficult" samples that satisfy the requirements of evaluating visual token compression methods.
The pipeline consists of three critical steps:
All inference results (raw data) can be downloaded in OneDrive.
Final evaluation results can be found in Final_Results.
conda create -n VTC python=3.10 -y
conda activate VTC
cd Qwen2-VL/transformers && pip install -e .
pip install accelerate qwen-vl-utils[decord]
pip install flash-attn --no-build-isolation
cd ../../lmms-eval && pip install -e .
pip install qwen-vl-utils
pip install flash-attention-softmax-n
bash scripts/dart.sh false [downsample_ratio]
bash scripts/dart.sh true 1 [reduction_ratio]
bash scripts/effivlm.sh 1 [reduction_ratio]
python tools/reorganize_data.py
Data list
βββ Qwen2-VL-7B-Instruct
βββ Downsample
βββ 1
π xxx.jsonl
βββ 2
βββ 3
βββ 4
βββ 5
βββ 10
βββ VisionZip
βββ 0.01
βββ 0.04
βββ 0.0625
βββ 0.1111
βββ 0.25
βββ PruMerge+
βββ FastV
βββ Llava-ov-7B
βββ Downsample
βββ VisionZip
βββ PruMerge+
βββ FastV
βββ DART
python tools/analyze_results.py --all
If you have any problems, please contact:
π§ cliao127@connect.hkust-gz.edu.cn
We will response and fix the problems ASAP! Thanks!
If you find this project helpful, please consider citing the following paper:
@article{liao2026usingrightbenchmarkevaluation,
title={Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods},
author={Liao, Chenfei and Wang, Wensong and Wen, Zichen and Zheng, Xu and Wang, Yiyu and He, Haocong and Lyu, Yuanhuiyi and Jiang, Lutao and Zou, Xin and Fu, Yuqian and Ren, Bin and Zhang, Linfeng and Hu, Xuming},
journal={arXiv preprint arXiv:2510.07143},
year={2026}
}
67 followers Β· starred Oct 2025
Python
89.3%
Shell
10.7%
[ACL2026 Main] Data & Code of "Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods"
Python
37
37 commits
updated Apr 9, 2026
Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
Chenfei Liao1,2,6 Wensong Wang3,2 Zichen Wen2,5 Xu Zheng1,4,6 Yiyu Wang2 Haocong He2
Yuanhuiyi Lyu1,6 Lutao Jiang1,6 Xin Zou1,6 Yuqian Fu4 Bin Ren7,8,4 Linfeng Zhang2,π§ Xuming Hu1,6,π§
1Hong Kong University of Science and Technology (Guangzhou) 2Shanghai Jiao Tong University
3Northeastern University 4INSAIT, Sofia University βSt. Kliment Ohridskiβ
5Shanghai AI Laboratory 6Hong Kong University of Science and Technology
7University of Pisa 8University of Trento
Recent efforts to accelerate inference in Multimodal Large Language Models (MLLMs) have largely focused on visual token compression. The effectiveness of these methods is commonly evaluated by measuring the accuracy drop on existing MLLM benchmarks before and after compression. However, these benchmarks are originally designed to assess general perception and reasoning abilities, rather than the specific challenges posed by visual token compression, leading to a fundamental task mismatch.
In this work, we uncover a counterintuitive yet consistent phenomenon: simple image downsampling outperforms many advanced visual token compression methods across multiple widely used benchmarks.
Through a comprehensive empirical study spanning eight popular benchmarks and multiple state-of-the-art compression techniques, we show that (i) current benchmarks contain substantial noise (task-irrelevant samples) for evaluating visual token compression, and (ii) downsampling can act as an effective data filter that distinguishes between simple and difficult samples with respect to compression sensitivity.
Motivated by these findings, we propose VTC-Bench, an evaluation framework that explicitly leverages downsampling as a discriminator to denoise existing benchmarks, enabling a fairer and more meaningful additional assessment of visual token compression methods.
Some recent MLLMs, such as Qwen2-VL and Qwen2.5-VL, natively support inputs of varying resolutions. A trivial yet efficient method to handle high-resolution images is to simply downsample them to a lower resolution. However, most token compression methods for MLLMs choose to adaptively drop useless tokens or merge similar tokens instead of directly downsampling the original image, which theoretically should be more intelligent.
Surprisingly, we find that image downsampling consistently exceeds other sophisticated methods under some settings. Based on comprehensive experiments, we propose a bold hypothesis:
Some data in the existing benchmarks is overly simplistic and irrelevant to evaluating visual token compression methods, leading to the unreasonable phenomenon that even the downsampling method is sufficient to deal with the visual token compression task.
To validate this, we design a data-centric analysis using downsampling as a discriminator. We identify two crucial findings:
Based on these findings, we propose VTC-Bench, a new evaluation framework specifically designed to optimize and denoise current existing benchmarks. By explicitly distinguishing between βsimpleβ and βdifficultβ samples through downsampling, VTC-Bench adaptively selects "difficult" samples that satisfy the requirements of evaluating visual token compression methods.
The pipeline consists of three critical steps:
All inference results (raw data) can be downloaded in OneDrive.
Final evaluation results can be found in Final_Results.
conda create -n VTC python=3.10 -y
conda activate VTC
cd Qwen2-VL/transformers && pip install -e .
pip install accelerate qwen-vl-utils[decord]
pip install flash-attn --no-build-isolation
cd ../../lmms-eval && pip install -e .
pip install qwen-vl-utils
pip install flash-attention-softmax-n
bash scripts/dart.sh false [downsample_ratio]
bash scripts/dart.sh true 1 [reduction_ratio]
bash scripts/effivlm.sh 1 [reduction_ratio]
python tools/reorganize_data.py
Data list
βββ Qwen2-VL-7B-Instruct
βββ Downsample
βββ 1
π xxx.jsonl
βββ 2
βββ 3
βββ 4
βββ 5
βββ 10
βββ VisionZip
βββ 0.01
βββ 0.04
βββ 0.0625
βββ 0.1111
βββ 0.25
βββ PruMerge+
βββ FastV
βββ Llava-ov-7B
βββ Downsample
βββ VisionZip
βββ PruMerge+
βββ FastV
βββ DART
python tools/analyze_results.py --all
If you have any problems, please contact:
π§ cliao127@connect.hkust-gz.edu.cn
We will response and fix the problems ASAP! Thanks!
If you find this project helpful, please consider citing the following paper:
@article{liao2026usingrightbenchmarkevaluation,
title={Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods},
author={Liao, Chenfei and Wang, Wensong and Wen, Zichen and Zheng, Xu and Wang, Yiyu and He, Haocong and Lyu, Yuanhuiyi and Jiang, Lutao and Zou, Xin and Fu, Yuqian and Ren, Bin and Zhang, Linfeng and Hu, Xuming},
journal={arXiv preprint arXiv:2510.07143},
year={2026}
}
67 followers Β· starred Oct 2025
Python
89.3%
Shell
10.7%