CapRL-Video-178K.jsonl Video Path Setup
9
5 commits
7 linked in READMEs
updated Jun 10, 2026
Each value is a relative path under the Hugging Face dataset root of lmms-lab/LLaVA-Video-178K.
Example:
"video": "0_30_s_academic_v0_1/videos/academic_source/activitynet/v_01vNlQLepsE.mp4"
Download the original videos from Hugging Face:
0_30_s_youtube_v0_1: 72970 samples2_3_m_youtube_v0_1: 24685 samples1_2_m_youtube_v0_1: 22427 samples30_60_s_youtube_v0_1: 19994 samples0_30_s_academic_v0_1: 12139 samples30_60_s_academic_v0_1: 10503 samples1_2_m_academic_v0_1: 4572 samples2_3_m_academic_v0_1: 3089 samplesThe videos in these folders are distributed on Hugging Face as *_videos_*.tar.gz archives, together with processed annotation JSON files. The annotation JSON files are not required for CapRL-Video-178K.jsonl; only the extracted video files are needed.
After downloading and extracting the archives, organize all split folders under one dataset root:
/path/to/LLaVA-Video-178K/
βββ 0_30_s_academic_v0_1/
β βββ videos/
β βββ academic_source/
βββ 0_30_s_youtube_v0_1/
β βββ videos/
β βββ liwei_youtube_videos/
βββ 1_2_m_academic_v0_1/
β βββ academic_source/
βββ 1_2_m_youtube_v0_1/
β βββ liwei_youtube_videos/
βββ 2_3_m_academic_v0_1/
β βββ academic_source/
βββ 2_3_m_youtube_v0_1/
β βββ liwei_youtube_videos/
βββ 30_60_s_academic_v0_1/
β βββ academic_source/
βββ 30_60_s_youtube_v0_1/
βββ liwei_youtube_videos/
The values in video should be joined with /path/to/LLaVA-Video-178K. For example:
from pathlib import Path
video_root = Path('/path/to/LLaVA-Video-178K')
video_path = video_root / sample['video']
huggingface-cli download lmms-lab/LLaVA-Video-178K \
--repo-type dataset \
--local-dir /path/to/LLaVA-Video-178K \
--include '0_30_s_academic_v0_1/*' '0_30_s_youtube_v0_1/*' '1_2_m_academic_v0_1/*' '1_2_m_youtube_v0_1/*' '2_3_m_academic_v0_1/*' '2_3_m_youtube_v0_1/*' '30_60_s_academic_v0_1/*' '30_60_s_youtube_v0_1/*'
cd /path/to/LLaVA-Video-178K
for f in 0_30_s_academic_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 0_30_s_academic_v0_1; done
for f in 0_30_s_youtube_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 0_30_s_youtube_v0_1; done
for f in 1_2_m_academic_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 1_2_m_academic_v0_1; done
for f in 1_2_m_youtube_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 1_2_m_youtube_v0_1; done
for f in 2_3_m_academic_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 2_3_m_academic_v0_1; done
for f in 2_3_m_youtube_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 2_3_m_youtube_v0_1; done
for f in 30_60_s_academic_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 30_60_s_academic_v0_1; done
for f in 30_60_s_youtube_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 30_60_s_youtube_v0_1; done
If your downloader places files in a different location, keep the extracted files under the same split-level relative paths shown above, or update your training script to join sample['video'] with your actual LLaVA-Video-178K root.
Long Xing* Β· Xiaoyi Dong* Β· Yuhang Zang Β· Yuhang Cao Β· Jianze Liang Β· Qidong Huang Β· Jiaqi Wang Β· Feng Wu Β· Dahua Lin
Penghui Yang* Β· Long Xing* Β· Xiaoyi Dong Β· Yuhang Zang Β· Yuhang Cao Β· Yibin Wang Β· Yujie Zhou Β· Jiazi Bu Β· Jianze Liang Β· Qidong Huang Β· Jiaqi Wang Β· Feng Wu Β· Dahua Lin
πCapRL++ Paper | πCapRL Paper | π Github | π€CapRL Collection| Series | Models & Resources |
|---|---|
| CapRL 3.0 Series (CapRL++) | π€ CapRL-Video-4B | π CapRL-Video-178K Dataset |
| CapRL 2.0 Series | π€ CapRL-Qwen3VL-2B | π€ CapRL-Qwen3VL-4B | π¦ CapRL-Qwen3VL-2B-GGUF | π¦ CapRL-Qwen3VL-4B-GGUF | πCapRL-Qwen3VL-4B Space |
| CapRL 1.0 Series | π€ CapRL-Qwen2.5VL-3B | π€ CapRL-InternVL3.5-8B | π CapRL-2M Dataset | π¦ CapRL-3B-GGUF | π¦ CapRL-3B-i1-GGUF | πCapRL-Qwen2.5VL-3B Space |
CapRL 3.0 series (CapRL++): CapRL-Video-4B has been released! CapRL++ extends the original image-caption RL framework to a unified image and video captioning paradigm with verifiable rewards.
We are excited to release the CapRL 2.0 series: CapRL-Qwen3VL-2B and CapRL-Qwen3VL-4B. These models feature fewer parameters while delivering even more powerful captioning performance. Notably, CapRL-Qwen3VL-2B outperforms both CapRL-Qwen2.5VL-3B and Qwen2.5VL-72B in captioning tasks, while CapRL-Qwen3VL-4B further demonstrates a significant performance leap over the 2B version. This improvement in efficiency is driven by our upgraded training recipe, which includes a more rigorous QA data filter and a significantly more diverse image dataset. We welcome everyone to try them out!
When selecting between the available CapRL models, it's essential to consider the trade-off between performance and computational cost. This guide will help you choose the most suitable model for your specific needs:
| Model | Parameters | Strength |
|---|---|---|
| π€CapRL-Qwen3VL-2B | 2B | Speed, Efficiency |
| π€CapRL-Qwen3VL-4B | 4B | High Performance, Advanced Captioning Ability |
| π€CapRL-Video-4B | 4B | Extremely Dense Video Captioning |
Now you can try out CapRL with your own imagesπ¨!Β Β Β Β β‘οΈΒ Β Β Β πCapRL-Qwen2.5VL-3B Space and πCapRL-Qwen3VL-4B Space.
We are working on even stronger base models and upgrading our training recipe β stay tuned!
CapRL++ folder.π We are excited to introduce the CapRL series, a family of dense captioning models trained with reinforcement learning rather than conventional supervised caption imitation.
The original CapRL framework focuses on dense image captioning. It optimizes an LVLM captioner with QA-derived rewards: a caption is considered high quality when a text-only model can answer visual questions using only that caption. With this recipe, the lightweight CapRL-3B achieves perception capabilities comparable to Qwen2.5-VL-72B.
CapRL++ further generalizes this idea from static images to dynamic videos. It trains a Qwen3-VL-based captioner with a unified RLVR pipeline, where generated captions are evaluated by their downstream utility for multiple-choice visual question answering. For videos, CapRL++ adds timestamp-format rewards and length-aware regularization so the model learns dense, temporally grounded, and non-redundant descriptions.
CapRL++ is the video-oriented extension of CapRL. It keeps the central principle of CapRL: a caption should be rewarded by how useful it is for downstream visual question answering. Instead of comparing a generated caption with a fixed reference, CapRL++ lets the policy model generate captions, then asks a separate vision-free LLM to answer curated multiple-choice questions using only those captions. The answer accuracy becomes a verifiable reward for RL training.
For a sampled caption c, CapRL++ uses a multidimensional reward:
R_total(c) = R_acc(c) + alpha * R_format(c) + beta * R_len(c)
R_acc): measures whether a text-only LLM can answer image/video MCQs from the generated caption alone. Options are shuffled and sampled multiple times to reduce answer-position bias.R_format): used for video captions. It encourages valid timestamp brackets and chronological ordering, helping the model produce temporally grounded narratives.R_len): discourages reward hacking through overly long or repetitive captions, pushing the model toward high information density.CapRL++ uses S2D-Boot, a two-stage image-to-video training recipe:
The CapRL++ implementation is in CapRL++:
CapRL++/
βββ train/
β βββ scripts/ # reward service and verl training launch scripts
β βββ verl/ # bundled verl backend with video caption RL recipe
βββ eval/
βββ scripts/ # Prism video evaluation scripts
βββ tools/ # benchmark judge helpers
βββ README.md
For details, see:
For CapRL image training and evaluation:
git clone https://github.com/InternLM/CapRL.git
cd CapRL/CapRL_Training
conda create -n CapRL python=3.10
conda activate CapRL
bash setup.sh
The setup.sh will sequentially:
pip install -e .
For CapRL++ video training and evaluation:cd CapRL/CapRL++/train
conda create -n caprl python=3.10 -y
conda activate caprl
pip install -r scripts/requirements.txt
pip install -e ./verl
Video Prism evaluation dependencies are installed separately:
cd CapRL/CapRL++/eval
pip install -r requirements.txt
If you want to use CapRL-3B for captioning, you can directly follow the exact same inference approach as in Qwen2.5-VL-series.
The prompt we use for training and evaluation is Please describe this image in detail.
We recommend using vLLM to speed up inference.
For CapRL-Video-4B, use the Qwen3-VL video inference interface or the Prism evaluation scripts under CapRL++/eval. A typical video caption prompt is:
Please describe this video in detail.
Run the command below to start an OpenAI-compatible API service:
vllm serve "/PATH/CapRL-3B" \
--trust-remote-code \
--tensor-parallel-size=1 \
--pipeline-parallel-size=1 \
--gpu_memory_utilization=0.95 \
--served-model-name=caprl \
--port 8000 \
--host 0.0.0.0
Then you can use the chat API as below: (see OpenAI API protocol document for more details):
import base64
from openai import OpenAI
# Set OpenAI's API key and API base to use vLLM's API server.
openai_api_key = "EMPTY"
openai_api_base = "http://localhost:8000/v1"
client = OpenAI(
api_key=openai_api_key,
base_url=openai_api_base,
)
image_path = "/path/to/local/image.png"
with open(image_path, "rb") as f:
encoded_image = base64.b64encode(f.read())
encoded_image_text = encoded_image.decode("utf-8")
base64_qwen = f"data:image;base64,{encoded_image_text}"
chat_response = client.chat.completions.create(
model="caprl",
messages=[
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": base64_qwen
},
},
{"type": "text", "text": "Please describe this image in detail."},
],
},
],
temperature=1.0,
max_tokens=max_tokens,
top_p=1.0,
extra_body={
"repetition_penalty": 1.0,
},
)
print("Chat response:", chat_response)
This part of the code is in the QA_data_curation folder, which contains all four steps for generating QA data:
ROTATE_NUM controls how many times each question is answered. If a question is answered only once, the randomness may be too high and can easily lead to misjudgment.visual acc higher than 0.75 and text acc lower than 0.25 to avoid data leakage and ensure the model can correctly answer questions when images are provided.All training scripts are located in CapRL_Training/scripts/. Taking qwen2.5vl3b_75k_reward_qwen2.5_3b as an example:
Step 1: Start the reward server
cd CapRL_Training
bash scripts/qwen2.5vl3b_75k_reward_qwen2.5_3b/reward/rjob.sh
Once the reward server is running, note its IP address.
Step 2: Launch training
Set <REWARD_SERVER_IP> in training/launch.sh to the IP from Step 1, then:
bash scripts/qwen2.5vl3b_75k_reward_qwen2.5_3b/training/rjob.sh
Note: The training scripts require
vllm>=0.11.0for Qwen3-VL compatibility. However, the reward server using Qwen2.5/Qwen3 LLM may occasionally encounter issues with higher vLLM versions. We recommend running the reward server in a separate conda environment with a lower version such asvllm==0.10.1.
A note on migrating CapRL to other codebases: Our training code is built on OpenRLHF, which originally lacked VLM (e.g., Qwen3-VL) RL training support. We added VLM adaptation and CapRL's two-stage reward on top of it. If you prefer a more lightweight alternative, consider using VeRL, which natively supports VLM training β you only need to customize the reward computation (e.g., by querying a vLLM reward server). If there is demand for VeRL integration, please open an issue to let us know.
Our CapRL-2M dataset is available on : π Hugging Face
It includes images from ShareGPT-1M and DenseFusion-1M, with high-quality captions re-annotated using CapRL-3B, totaling 2M samples.
In our JSONL files, we provide the captions along with their corresponding image paths. The images can be downloaded from ShareGPT-1M and DenseFusion-1M.
To reproduce the pretraining experiments presented in our paper:
Initialize Qwen2.5-VL.
Follow the steps in the notebook initiallize_vlm_3b.ipynb to set up the Qwen2.5-VL model for training.
Training.
We use LLaMA-Factory for pretraining. The training scripts are provided in Pretraining_exp/scripts/, covering all 3 stages:
Stage0_initial_align.sh β Initial alignment with LLaVA-558KStage1_further_pretrain.sh β Further pretraining with CapRL-1M caption dataStage2_sft.sh β SFT with general instruction data, Open-LLaVA-NeXT-1MWe evaluate caption quality by decoupling the traditional VQA (Visual Question Answering) task:
This approach allows us to assess the informational quality and completeness of the generated captions β if the language model can accurately answer visual questions based only on the caption, then the caption is likely high-quality.
The complete evaluation scripts can be found in the Prism_Evaluation folder, with the core implementation located in Eval_CapRL.py.
The Prism evaluation files are available at CapRL-Evaluation-Files. The dataset contains json_file/ for the evaluation JSON files and bench_image_folder.zip for the corresponding images.
huggingface-cli download internlm/CapRL-Evaluation-Files --repo-type dataset --local-dir CapRL-Evaluation-Files
cd CapRL-Evaluation-Files
unzip bench_image_folder.zip
Use the JSON files under json_file/ as --data-path and pass the dataset root as --image-root. The image paths inside each JSON are relative to the dataset root, for example bench_image_folder/lmm_eval_chartqa/41699051005347.png.
python -m Eval_CapRL \
--data-path /path/to/CapRL-Evaluation-Files/json_file/lmm_eval_chartqa.json \
--image-root /path/to/CapRL-Evaluation-Files \
--tag chartqa \
...
The model used for answering questions based on captions is CapRL-Eval-3B, which is a finetuned version of Qwen2.5-VL-3B. When dealing with tasks such as ChartQA (not multiple-choice questions), it provides more stable output formatting.
You can specify --reward-model-path as the path to CapRL-Eval-3B in Eval_CapRL.py.
Usage and License Notices: The data and code are intended and licensed for research use only. License: Attribution-NonCommercial 4.0 International It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
If you find CapRL++ useful for your research, please consider citing:
@article{yang2026caprlplusplus,
title={CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning},
author={Yang, Penghui and Xing, Long and Dong, Xiaoyi and Zang, Yuhang and Cao, Yuhang and Wang, Yibin and Zhou, Yujie and Bu, Jiazi and Liang, Jianze and Huang, Qidong and Wang, Jiaqi and Wu, Feng and Lin, Dahua},
journal={arXiv preprint arXiv:2606.09393},
year={2026}
}
For the original CapRL paper:
@article{xing2025caprl,
title={CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning},
author={Xing, Long and Dong, Xiaoyi and Zang, Yuhang and Cao, Yuhang and Liang, Jianze and Huang, Qidong and Wang, Jiaqi and Wu, Feng and Lin, Dahua},
journal={arXiv preprint arXiv:2509.22647},
year={2025}
}
CapRL-Video-178K.jsonl Video Path Setup
9
5 commits
7 linked in READMEs
updated Jun 10, 2026
Each value is a relative path under the Hugging Face dataset root of lmms-lab/LLaVA-Video-178K.
Example:
"video": "0_30_s_academic_v0_1/videos/academic_source/activitynet/v_01vNlQLepsE.mp4"
Download the original videos from Hugging Face:
0_30_s_youtube_v0_1: 72970 samples2_3_m_youtube_v0_1: 24685 samples1_2_m_youtube_v0_1: 22427 samples30_60_s_youtube_v0_1: 19994 samples0_30_s_academic_v0_1: 12139 samples30_60_s_academic_v0_1: 10503 samples1_2_m_academic_v0_1: 4572 samples2_3_m_academic_v0_1: 3089 samplesThe videos in these folders are distributed on Hugging Face as *_videos_*.tar.gz archives, together with processed annotation JSON files. The annotation JSON files are not required for CapRL-Video-178K.jsonl; only the extracted video files are needed.
After downloading and extracting the archives, organize all split folders under one dataset root:
/path/to/LLaVA-Video-178K/
βββ 0_30_s_academic_v0_1/
β βββ videos/
β βββ academic_source/
βββ 0_30_s_youtube_v0_1/
β βββ videos/
β βββ liwei_youtube_videos/
βββ 1_2_m_academic_v0_1/
β βββ academic_source/
βββ 1_2_m_youtube_v0_1/
β βββ liwei_youtube_videos/
βββ 2_3_m_academic_v0_1/
β βββ academic_source/
βββ 2_3_m_youtube_v0_1/
β βββ liwei_youtube_videos/
βββ 30_60_s_academic_v0_1/
β βββ academic_source/
βββ 30_60_s_youtube_v0_1/
βββ liwei_youtube_videos/
The values in video should be joined with /path/to/LLaVA-Video-178K. For example:
from pathlib import Path
video_root = Path('/path/to/LLaVA-Video-178K')
video_path = video_root / sample['video']
huggingface-cli download lmms-lab/LLaVA-Video-178K \
--repo-type dataset \
--local-dir /path/to/LLaVA-Video-178K \
--include '0_30_s_academic_v0_1/*' '0_30_s_youtube_v0_1/*' '1_2_m_academic_v0_1/*' '1_2_m_youtube_v0_1/*' '2_3_m_academic_v0_1/*' '2_3_m_youtube_v0_1/*' '30_60_s_academic_v0_1/*' '30_60_s_youtube_v0_1/*'
cd /path/to/LLaVA-Video-178K
for f in 0_30_s_academic_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 0_30_s_academic_v0_1; done
for f in 0_30_s_youtube_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 0_30_s_youtube_v0_1; done
for f in 1_2_m_academic_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 1_2_m_academic_v0_1; done
for f in 1_2_m_youtube_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 1_2_m_youtube_v0_1; done
for f in 2_3_m_academic_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 2_3_m_academic_v0_1; done
for f in 2_3_m_youtube_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 2_3_m_youtube_v0_1; done
for f in 30_60_s_academic_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 30_60_s_academic_v0_1; done
for f in 30_60_s_youtube_v0_1/*_videos_*.tar.gz; do [ -e "$f" ] && tar -xzf "$f" -C 30_60_s_youtube_v0_1; done
If your downloader places files in a different location, keep the extracted files under the same split-level relative paths shown above, or update your training script to join sample['video'] with your actual LLaVA-Video-178K root.
Long Xing* Β· Xiaoyi Dong* Β· Yuhang Zang Β· Yuhang Cao Β· Jianze Liang Β· Qidong Huang Β· Jiaqi Wang Β· Feng Wu Β· Dahua Lin
Penghui Yang* Β· Long Xing* Β· Xiaoyi Dong Β· Yuhang Zang Β· Yuhang Cao Β· Yibin Wang Β· Yujie Zhou Β· Jiazi Bu Β· Jianze Liang Β· Qidong Huang Β· Jiaqi Wang Β· Feng Wu Β· Dahua Lin
πCapRL++ Paper | πCapRL Paper | π Github | π€CapRL Collection| Series | Models & Resources |
|---|---|
| CapRL 3.0 Series (CapRL++) | π€ CapRL-Video-4B | π CapRL-Video-178K Dataset |
| CapRL 2.0 Series | π€ CapRL-Qwen3VL-2B | π€ CapRL-Qwen3VL-4B | π¦ CapRL-Qwen3VL-2B-GGUF | π¦ CapRL-Qwen3VL-4B-GGUF | πCapRL-Qwen3VL-4B Space |
| CapRL 1.0 Series | π€ CapRL-Qwen2.5VL-3B | π€ CapRL-InternVL3.5-8B | π CapRL-2M Dataset | π¦ CapRL-3B-GGUF | π¦ CapRL-3B-i1-GGUF | πCapRL-Qwen2.5VL-3B Space |
CapRL 3.0 series (CapRL++): CapRL-Video-4B has been released! CapRL++ extends the original image-caption RL framework to a unified image and video captioning paradigm with verifiable rewards.
We are excited to release the CapRL 2.0 series: CapRL-Qwen3VL-2B and CapRL-Qwen3VL-4B. These models feature fewer parameters while delivering even more powerful captioning performance. Notably, CapRL-Qwen3VL-2B outperforms both CapRL-Qwen2.5VL-3B and Qwen2.5VL-72B in captioning tasks, while CapRL-Qwen3VL-4B further demonstrates a significant performance leap over the 2B version. This improvement in efficiency is driven by our upgraded training recipe, which includes a more rigorous QA data filter and a significantly more diverse image dataset. We welcome everyone to try them out!
When selecting between the available CapRL models, it's essential to consider the trade-off between performance and computational cost. This guide will help you choose the most suitable model for your specific needs:
| Model | Parameters | Strength |
|---|---|---|
| π€CapRL-Qwen3VL-2B | 2B | Speed, Efficiency |
| π€CapRL-Qwen3VL-4B | 4B | High Performance, Advanced Captioning Ability |
| π€CapRL-Video-4B | 4B | Extremely Dense Video Captioning |
Now you can try out CapRL with your own imagesπ¨!Β Β Β Β β‘οΈΒ Β Β Β πCapRL-Qwen2.5VL-3B Space and πCapRL-Qwen3VL-4B Space.
We are working on even stronger base models and upgrading our training recipe β stay tuned!
CapRL++ folder.π We are excited to introduce the CapRL series, a family of dense captioning models trained with reinforcement learning rather than conventional supervised caption imitation.
The original CapRL framework focuses on dense image captioning. It optimizes an LVLM captioner with QA-derived rewards: a caption is considered high quality when a text-only model can answer visual questions using only that caption. With this recipe, the lightweight CapRL-3B achieves perception capabilities comparable to Qwen2.5-VL-72B.
CapRL++ further generalizes this idea from static images to dynamic videos. It trains a Qwen3-VL-based captioner with a unified RLVR pipeline, where generated captions are evaluated by their downstream utility for multiple-choice visual question answering. For videos, CapRL++ adds timestamp-format rewards and length-aware regularization so the model learns dense, temporally grounded, and non-redundant descriptions.
CapRL++ is the video-oriented extension of CapRL. It keeps the central principle of CapRL: a caption should be rewarded by how useful it is for downstream visual question answering. Instead of comparing a generated caption with a fixed reference, CapRL++ lets the policy model generate captions, then asks a separate vision-free LLM to answer curated multiple-choice questions using only those captions. The answer accuracy becomes a verifiable reward for RL training.
For a sampled caption c, CapRL++ uses a multidimensional reward:
R_total(c) = R_acc(c) + alpha * R_format(c) + beta * R_len(c)
R_acc): measures whether a text-only LLM can answer image/video MCQs from the generated caption alone. Options are shuffled and sampled multiple times to reduce answer-position bias.R_format): used for video captions. It encourages valid timestamp brackets and chronological ordering, helping the model produce temporally grounded narratives.R_len): discourages reward hacking through overly long or repetitive captions, pushing the model toward high information density.CapRL++ uses S2D-Boot, a two-stage image-to-video training recipe:
The CapRL++ implementation is in CapRL++:
CapRL++/
βββ train/
β βββ scripts/ # reward service and verl training launch scripts
β βββ verl/ # bundled verl backend with video caption RL recipe
βββ eval/
βββ scripts/ # Prism video evaluation scripts
βββ tools/ # benchmark judge helpers
βββ README.md
For details, see:
For CapRL image training and evaluation:
git clone https://github.com/InternLM/CapRL.git
cd CapRL/CapRL_Training
conda create -n CapRL python=3.10
conda activate CapRL
bash setup.sh
The setup.sh will sequentially:
pip install -e .
For CapRL++ video training and evaluation:cd CapRL/CapRL++/train
conda create -n caprl python=3.10 -y
conda activate caprl
pip install -r scripts/requirements.txt
pip install -e ./verl
Video Prism evaluation dependencies are installed separately:
cd CapRL/CapRL++/eval
pip install -r requirements.txt
If you want to use CapRL-3B for captioning, you can directly follow the exact same inference approach as in Qwen2.5-VL-series.
The prompt we use for training and evaluation is Please describe this image in detail.
We recommend using vLLM to speed up inference.
For CapRL-Video-4B, use the Qwen3-VL video inference interface or the Prism evaluation scripts under CapRL++/eval. A typical video caption prompt is:
Please describe this video in detail.
Run the command below to start an OpenAI-compatible API service:
vllm serve "/PATH/CapRL-3B" \
--trust-remote-code \
--tensor-parallel-size=1 \
--pipeline-parallel-size=1 \
--gpu_memory_utilization=0.95 \
--served-model-name=caprl \
--port 8000 \
--host 0.0.0.0
Then you can use the chat API as below: (see OpenAI API protocol document for more details):
import base64
from openai import OpenAI
# Set OpenAI's API key and API base to use vLLM's API server.
openai_api_key = "EMPTY"
openai_api_base = "http://localhost:8000/v1"
client = OpenAI(
api_key=openai_api_key,
base_url=openai_api_base,
)
image_path = "/path/to/local/image.png"
with open(image_path, "rb") as f:
encoded_image = base64.b64encode(f.read())
encoded_image_text = encoded_image.decode("utf-8")
base64_qwen = f"data:image;base64,{encoded_image_text}"
chat_response = client.chat.completions.create(
model="caprl",
messages=[
{
"role": "user",
"content": [
{
"type": "image_url",
"image_url": {
"url": base64_qwen
},
},
{"type": "text", "text": "Please describe this image in detail."},
],
},
],
temperature=1.0,
max_tokens=max_tokens,
top_p=1.0,
extra_body={
"repetition_penalty": 1.0,
},
)
print("Chat response:", chat_response)
This part of the code is in the QA_data_curation folder, which contains all four steps for generating QA data:
ROTATE_NUM controls how many times each question is answered. If a question is answered only once, the randomness may be too high and can easily lead to misjudgment.visual acc higher than 0.75 and text acc lower than 0.25 to avoid data leakage and ensure the model can correctly answer questions when images are provided.All training scripts are located in CapRL_Training/scripts/. Taking qwen2.5vl3b_75k_reward_qwen2.5_3b as an example:
Step 1: Start the reward server
cd CapRL_Training
bash scripts/qwen2.5vl3b_75k_reward_qwen2.5_3b/reward/rjob.sh
Once the reward server is running, note its IP address.
Step 2: Launch training
Set <REWARD_SERVER_IP> in training/launch.sh to the IP from Step 1, then:
bash scripts/qwen2.5vl3b_75k_reward_qwen2.5_3b/training/rjob.sh
Note: The training scripts require
vllm>=0.11.0for Qwen3-VL compatibility. However, the reward server using Qwen2.5/Qwen3 LLM may occasionally encounter issues with higher vLLM versions. We recommend running the reward server in a separate conda environment with a lower version such asvllm==0.10.1.
A note on migrating CapRL to other codebases: Our training code is built on OpenRLHF, which originally lacked VLM (e.g., Qwen3-VL) RL training support. We added VLM adaptation and CapRL's two-stage reward on top of it. If you prefer a more lightweight alternative, consider using VeRL, which natively supports VLM training β you only need to customize the reward computation (e.g., by querying a vLLM reward server). If there is demand for VeRL integration, please open an issue to let us know.
Our CapRL-2M dataset is available on : π Hugging Face
It includes images from ShareGPT-1M and DenseFusion-1M, with high-quality captions re-annotated using CapRL-3B, totaling 2M samples.
In our JSONL files, we provide the captions along with their corresponding image paths. The images can be downloaded from ShareGPT-1M and DenseFusion-1M.
To reproduce the pretraining experiments presented in our paper:
Initialize Qwen2.5-VL.
Follow the steps in the notebook initiallize_vlm_3b.ipynb to set up the Qwen2.5-VL model for training.
Training.
We use LLaMA-Factory for pretraining. The training scripts are provided in Pretraining_exp/scripts/, covering all 3 stages:
Stage0_initial_align.sh β Initial alignment with LLaVA-558KStage1_further_pretrain.sh β Further pretraining with CapRL-1M caption dataStage2_sft.sh β SFT with general instruction data, Open-LLaVA-NeXT-1MWe evaluate caption quality by decoupling the traditional VQA (Visual Question Answering) task:
This approach allows us to assess the informational quality and completeness of the generated captions β if the language model can accurately answer visual questions based only on the caption, then the caption is likely high-quality.
The complete evaluation scripts can be found in the Prism_Evaluation folder, with the core implementation located in Eval_CapRL.py.
The Prism evaluation files are available at CapRL-Evaluation-Files. The dataset contains json_file/ for the evaluation JSON files and bench_image_folder.zip for the corresponding images.
huggingface-cli download internlm/CapRL-Evaluation-Files --repo-type dataset --local-dir CapRL-Evaluation-Files
cd CapRL-Evaluation-Files
unzip bench_image_folder.zip
Use the JSON files under json_file/ as --data-path and pass the dataset root as --image-root. The image paths inside each JSON are relative to the dataset root, for example bench_image_folder/lmm_eval_chartqa/41699051005347.png.
python -m Eval_CapRL \
--data-path /path/to/CapRL-Evaluation-Files/json_file/lmm_eval_chartqa.json \
--image-root /path/to/CapRL-Evaluation-Files \
--tag chartqa \
...
The model used for answering questions based on captions is CapRL-Eval-3B, which is a finetuned version of Qwen2.5-VL-3B. When dealing with tasks such as ChartQA (not multiple-choice questions), it provides more stable output formatting.
You can specify --reward-model-path as the path to CapRL-Eval-3B in Eval_CapRL.py.
Usage and License Notices: The data and code are intended and licensed for research use only. License: Attribution-NonCommercial 4.0 International It should abide by the policy of OpenAI: https://openai.com/policies/terms-of-use
If you find CapRL++ useful for your research, please consider citing:
@article{yang2026caprlplusplus,
title={CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning},
author={Yang, Penghui and Xing, Long and Dong, Xiaoyi and Zang, Yuhang and Cao, Yuhang and Wang, Yibin and Zhou, Yujie and Bu, Jiazi and Liang, Jianze and Huang, Qidong and Wang, Jiaqi and Wu, Feng and Lin, Dahua},
journal={arXiv preprint arXiv:2606.09393},
year={2026}
}
For the original CapRL paper:
@article{xing2025caprl,
title={CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning},
author={Xing, Long and Dong, Xiaoyi and Zang, Yuhang and Cao, Yuhang and Liang, Jianze and Huang, Qidong and Wang, Jiaqi and Wu, Feng and Lin, Dahua},
journal={arXiv preprint arXiv:2509.22647},
year={2025}
}