This Space demonstrates CapRL-Qwen3VL-4B, a 4B parameter multimodal vision-language model fine-tuned using reinforcement learning for dense image captioning.
Upload an image or select from the examples to generate a detailed caption.
@article{xing2025caprl,
title={{CapRL}: Stimulating Dense Image Caption Capabilities via Reinforcement Learning},
author={Xing, Long and Dong, Xiaoyi and Zang, Yuhang and Cao, Yuhang and Liang, Jianze and Huang, Qidong and Wang, Jiaqi and Wu, Feng and Lin, Dahua},
journal={arXiv preprint arXiv:2509.22647},
year={2025}
}
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference
This Space demonstrates CapRL-Qwen3VL-4B, a 4B parameter multimodal vision-language model fine-tuned using reinforcement learning for dense image captioning.
Upload an image or select from the examples to generate a detailed caption.
@article{xing2025caprl,
title={{CapRL}: Stimulating Dense Image Caption Capabilities via Reinforcement Learning},
author={Xing, Long and Dong, Xiaoyi and Zang, Yuhang and Cao, Yuhang and Liang, Jianze and Huang, Qidong and Wang, Jiaqi and Wu, Feng and Lin, Dahua},
journal={arXiv preprint arXiv:2509.22647},
year={2025}
}
Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference