yuhangzang/CapRL-Qwen3VL-4B

Space

CapRL-Qwen3VL-4B Image Captioning

6

5 commits

1 linked in READMEs

updated Dec 26, 2025

See the code

README

CapRL-Qwen3VL-4B Image Captioning

This Space demonstrates CapRL-Qwen3VL-4B, a 4B parameter multimodal vision-language model fine-tuned using reinforcement learning for dense image captioning.

Model

Usage

Upload an image or select from the examples to generate a detailed caption.

Citation

@article{xing2025caprl,
  title={{CapRL}: Stimulating Dense Image Caption Capabilities via Reinforcement Learning},
  author={Xing, Long and Dong, Xiaoyi and Zang, Yuhang and Cao, Yuhang and Liang, Jianze and Huang, Qidong and Wang, Jiaqi and Wu, Feng and Lin, Dahua},
  journal={arXiv preprint arXiv:2509.22647},
  year={2025}
}

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference

gradio

yuhangzang/CapRL-Qwen3VL-4B

Space

CapRL-Qwen3VL-4B Image Captioning

6

5 commits

1 linked in READMEs

updated Dec 26, 2025

See the code

README

CapRL-Qwen3VL-4B Image Captioning

This Space demonstrates CapRL-Qwen3VL-4B, a 4B parameter multimodal vision-language model fine-tuned using reinforcement learning for dense image captioning.

Model

Usage

Upload an image or select from the examples to generate a detailed caption.

Citation

@article{xing2025caprl,
  title={{CapRL}: Stimulating Dense Image Caption Capabilities via Reinforcement Learning},
  author={Xing, Long and Dong, Xiaoyi and Zang, Yuhang and Cao, Yuhang and Liang, Jianze and Huang, Qidong and Wang, Jiaqi and Wu, Feng and Lin, Dahua},
  journal={arXiv preprint arXiv:2509.22647},
  year={2025}
}

Check out the configuration reference at https://huggingface.co/docs/hub/spaces-config-reference

gradio