(ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’
Python
60
17 commits
updated Jan 26, 2026
Long Xing* · Qidong Huang* · Xiaoyi Dong · Pan Zhang · Yuhang Zang · Yuhang Cao · Jinsong Li · Shuangrui Ding · Weiming Zhang · Nenghai Yu · Jiaqi Wang · Feng Wu · Dahua Lin
📖Paper | 🤗Datasets | 🤗Daily Paper
🌈We introduce ScaleCap, an inference-time scalable image captioning strategy that generates comprehensive and detailed image captions. With ScaleCap, we construct a dataset containing 450k image-caption pairs for use by the open-source community. Our key observations highlight two inherent biases in LVLMs: multimodal bias resulting in imbalanced descriptive granularity; linguistic bias leading to hallucinated descriptions of non-existent objects. To address these issues, we propose two novel components: heuristic question answering and contrastive sentence rating. Extensive experiments demonstrate the effectiveness of ScaleCap.
git clone https://github.com/Cooperx521/ScaleCap.git
conda create -n ScaleCap python=3.10
conda activate ScaleCap
bash setup.sh
To quickly get started with generating captions using ScaleCap, we provide an example script. Simply run the following command:
bash scripts/launch_example.sh
In our setup, Qwen and its VL series are deployed using the vLLM framework.
We have verified that this setup works on NVIDIA A100 GPUs.
If your GPU has limited memory, we recommend doubling the number of devices specified by CUDA_VISIBLE_DEVICES to avoid out-of-memory issues.
You may need to modify the following lines in the script to better suit your hardware configuration:
Our ScaleCap450k dataset is available on : 🔗 Hugging Face
This dataset contains 450,000 images along with their corresponding captions generated by ScaleCap.
To reproduce the pretraining experiments presented in our paper:
Initialize Qwen2.5-VL.
Follow the steps in the notebook initiallize_vlm_3b.ipynb to set up the Qwen2.5-VL model for training.
Training. You can then use LLaMAFactory directly to run the training process.
We evaluate caption quality by decoupling the traditional VQA (Visual Question Answering) task:
This approach allows us to assess the informational quality and completeness of the generated captions — if the language model can accurately answer visual questions based only on the caption, then the caption is likely high-quality.
The inference pipeline can be found in the prism_benchmark directory.
The scripts used to compute evaluation metrics are located in the eval directory.
Usage and License Notices: The data and code are intended and licensed for research use only. License: Attribution-NonCommercial 4.0 International It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
Python
85.0%
Shell
7.6%
Jupyter Notebook
7.4%
(ICLR 2026)Official repository of 'ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing’
Python
60
17 commits
updated Jan 26, 2026
Long Xing* · Qidong Huang* · Xiaoyi Dong · Pan Zhang · Yuhang Zang · Yuhang Cao · Jinsong Li · Shuangrui Ding · Weiming Zhang · Nenghai Yu · Jiaqi Wang · Feng Wu · Dahua Lin
📖Paper | 🤗Datasets | 🤗Daily Paper
🌈We introduce ScaleCap, an inference-time scalable image captioning strategy that generates comprehensive and detailed image captions. With ScaleCap, we construct a dataset containing 450k image-caption pairs for use by the open-source community. Our key observations highlight two inherent biases in LVLMs: multimodal bias resulting in imbalanced descriptive granularity; linguistic bias leading to hallucinated descriptions of non-existent objects. To address these issues, we propose two novel components: heuristic question answering and contrastive sentence rating. Extensive experiments demonstrate the effectiveness of ScaleCap.
git clone https://github.com/Cooperx521/ScaleCap.git
conda create -n ScaleCap python=3.10
conda activate ScaleCap
bash setup.sh
To quickly get started with generating captions using ScaleCap, we provide an example script. Simply run the following command:
bash scripts/launch_example.sh
In our setup, Qwen and its VL series are deployed using the vLLM framework.
We have verified that this setup works on NVIDIA A100 GPUs.
If your GPU has limited memory, we recommend doubling the number of devices specified by CUDA_VISIBLE_DEVICES to avoid out-of-memory issues.
You may need to modify the following lines in the script to better suit your hardware configuration:
Our ScaleCap450k dataset is available on : 🔗 Hugging Face
This dataset contains 450,000 images along with their corresponding captions generated by ScaleCap.
To reproduce the pretraining experiments presented in our paper:
Initialize Qwen2.5-VL.
Follow the steps in the notebook initiallize_vlm_3b.ipynb to set up the Qwen2.5-VL model for training.
Training. You can then use LLaMAFactory directly to run the training process.
We evaluate caption quality by decoupling the traditional VQA (Visual Question Answering) task:
This approach allows us to assess the informational quality and completeness of the generated captions — if the language model can accurately answer visual questions based only on the caption, then the caption is likely high-quality.
The inference pipeline can be found in the prism_benchmark directory.
The scripts used to compute evaluation metrics are located in the eval directory.
Usage and License Notices: The data and code are intended and licensed for research use only. License: Attribution-NonCommercial 4.0 International It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
Python
85.0%
Shell
7.6%
Jupyter Notebook
7.4%