š¬ OraRL ā Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.
See the codeEfficient and scalable reinforcement learning for unified video MLLMs
Yunheng Li Ā· Guohong Mu Ā· Hao Li Ā· Shengsheng Qian Ā· Dingwen Zhang Ā· Qibin Hou Ā· Ming-Ming Cheng
š Paper Ā Ā Ā·Ā Ā š Project Page Ā Ā Ā·Ā Ā š® Live Demo Ā Ā Ā·Ā Ā š¤ Models (4B / 9B)
āļø Environment Ā Ā Ā·Ā Ā š Training Ā Ā Ā·Ā Ā š Evaluation Ā Ā Ā·Ā Ā āļø License
View all 11 verified evaluations on Papers with Code
ā¶ Click the image to watch the 1:38 project overview.
An OraRL update separates reliable annotation guidance from on-policy normalization:
This design uses task-native annotations directly and requires no chain-of-thought supervision or decoding.
Video-ORA-9B leads the matched seven-family comparison without CoT decoding.
Best and second-best values are highlighted per row; ā denotes an
original-report value whose frame, prompt, split, or decoding settings may
differ. Averages require complete family coverage.
| Model | Backbone | Released recipe | Weights |
|---|---|---|---|
| Video-ORA-9B | Qwen3.5-9B | orarl_9b.yaml | Hugging Face |
| Video-ORA-4B | Qwen3.5-4B | orarl_4b.yaml | Hugging Face |
Both Video-ORA checkpoints load directly with vLLM 0.19.1 for OpenAI-compatible serving:
MODEL=OraRL/Video-ORA-9B
vllm serve "$MODEL" \
--served-model-name Video-ORA-9B \
--trust-remote-code \
--dtype bfloat16 \
--tensor-parallel-size 1 \
--max-model-len 131072 \
--limit-mm-per-prompt '{"image": 1, "video": 1}'
Set --tensor-parallel-size to the GPU count for multi-GPU deployment and
lower --max-model-len on smaller-memory devices. Use
enable_thinking=false in the chat template for answer-only inference.
The release is organized around three user-facing workflows:
Training and evaluation are dry runs by default; inspect the resolved command
before adding --run. Checkpoints and evaluation media are hosted under the
OraRL Hugging Face organization.
OraRL is built on veRL ā a high-performance RL framework with HybridEngine. We thank its authors and contributors for open-sourcing the training infrastructure.
OraRL source is released under Apache-2.0. Datasets, models, benchmarks, and optional dependencies retain their original licenses; see NOTICE.
If you find OraRL useful, please consider giving this repository a ā and citing our paper.
@article{li2026orarl,
title = {Annotations as Rollouts: Efficient and Scalable
Reinforcement Learning for Video MLLMs},
author = {Li, Yunheng and Mu, Guohong and Li, Hao and
Qian, Shengsheng and Zhang, Dingwen and Hou, Qibin
and Cheng, Ming-Ming},
journal = {arXiv preprint arXiv:2608.20492},
year = {2026},
url = {https://arxiv.org/abs/2608.20492}
}
Python
92.9%
Shell
7.1%
š¬ OraRL ā Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.
See the codeEfficient and scalable reinforcement learning for unified video MLLMs
Yunheng Li Ā· Guohong Mu Ā· Hao Li Ā· Shengsheng Qian Ā· Dingwen Zhang Ā· Qibin Hou Ā· Ming-Ming Cheng
š Paper Ā Ā Ā·Ā Ā š Project Page Ā Ā Ā·Ā Ā š® Live Demo Ā Ā Ā·Ā Ā š¤ Models (4B / 9B)
āļø Environment Ā Ā Ā·Ā Ā š Training Ā Ā Ā·Ā Ā š Evaluation Ā Ā Ā·Ā Ā āļø License
View all 11 verified evaluations on Papers with Code
ā¶ Click the image to watch the 1:38 project overview.
An OraRL update separates reliable annotation guidance from on-policy normalization:
This design uses task-native annotations directly and requires no chain-of-thought supervision or decoding.
Video-ORA-9B leads the matched seven-family comparison without CoT decoding.
Best and second-best values are highlighted per row; ā denotes an
original-report value whose frame, prompt, split, or decoding settings may
differ. Averages require complete family coverage.
| Model | Backbone | Released recipe | Weights |
|---|---|---|---|
| Video-ORA-9B | Qwen3.5-9B | orarl_9b.yaml | Hugging Face |
| Video-ORA-4B | Qwen3.5-4B | orarl_4b.yaml | Hugging Face |
Both Video-ORA checkpoints load directly with vLLM 0.19.1 for OpenAI-compatible serving:
MODEL=OraRL/Video-ORA-9B
vllm serve "$MODEL" \
--served-model-name Video-ORA-9B \
--trust-remote-code \
--dtype bfloat16 \
--tensor-parallel-size 1 \
--max-model-len 131072 \
--limit-mm-per-prompt '{"image": 1, "video": 1}'
Set --tensor-parallel-size to the GPU count for multi-GPU deployment and
lower --max-model-len on smaller-memory devices. Use
enable_thinking=false in the chat template for answer-only inference.
The release is organized around three user-facing workflows:
Training and evaluation are dry runs by default; inspect the resolved command
before adding --run. Checkpoints and evaluation media are hosted under the
OraRL Hugging Face organization.
OraRL is built on veRL ā a high-performance RL framework with HybridEngine. We thank its authors and contributors for open-sourcing the training infrastructure.
OraRL source is released under Apache-2.0. Datasets, models, benchmarks, and optional dependencies retain their original licenses; see NOTICE.
If you find OraRL useful, please consider giving this repository a ā and citing our paper.
@article{li2026orarl,
title = {Annotations as Rollouts: Efficient and Scalable
Reinforcement Learning for Video MLLMs},
author = {Li, Yunheng and Mu, Guohong and Li, Hao and
Qian, Shengsheng and Zhang, Dingwen and Hou, Qibin
and Cheng, Ming-Ming},
journal = {arXiv preprint arXiv:2608.20492},
year = {2026},
url = {https://arxiv.org/abs/2608.20492}
}
Python
92.9%
Shell
7.1%