HVision-NKU/OraRL

šŸŽ¬ OraRL — Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.

Python

163

6 commits

updated Aug 28, 2026

See the code

README

简体中文

OraRL

Annotations as Rollouts

Efficient and scalable reinforcement learning for unified video MLLMs

Yunheng Li Ā· Guohong Mu Ā· Hao Li Ā· Shengsheng Qian Ā· Dingwen Zhang Ā· Qibin Hou Ā· Ming-Ming Cheng

šŸ“„ Paper Ā Ā Ā·Ā Ā  🌐 Project Page Ā Ā Ā·Ā Ā  šŸŽ® Live Demo Ā Ā Ā·Ā Ā  šŸ¤— Models (4B / 9B)

āš™ļø Environment Ā Ā Ā·Ā Ā  šŸš€ Training Ā Ā Ā·Ā Ā  šŸ“Š Evaluation Ā Ā Ā·Ā Ā  āš–ļø License

Papers with Code: SOTA on ActivityNet-TimeLens Papers with Code: SOTA on Charades-TimeLens Papers with Code: SOTA on MeViS Papers with Code: SOTA on QVHighlights-TimeLens Validation Papers with Code: SOTA on VideoHolmes Papers with Code: #2 on ReasonVOS

View all 11 verified evaluations on Papers with Code

Animated OraRL method preview

ā–¶ Click the image to watch the 1:38 project overview.

Why OraRL

  • Annotation-as-rollout: annotations become reliable positive rollouts while policy samples retain an on-policy baseline.
  • Seven task families: one update rule covers temporal and spatial grounding, segmentation, tracking, spatial-temporal grounding, video QA, and spatial intelligence.
  • Efficient training (4B): sign-balanced pruning delivers 1.48Ɨ faster updates (92.5 → 62.4 s/step) while reducing peak per-GPU memory from 62.4 to 50.9 GB.
  • Efficient inference: on one H20 with vLLM in BF16, weight loading occupies 8.6 GiB (4B) and 17.6 GiB (9B). On ten-minute, 2-fps videos, answer-only decoding cuts median post-TTFT latency from 4.78 s to 130 ms and total latency from 29.03 to 24.30 s.
  • Multimodal veRL infrastructure: a unified video contract carries cached artifacts, raw paths, or inline frame tensors through vLLM rollouts and FSDP updates, with decode-once frame reuse, temporal metadata, task-grouped batching, asynchronous Ray rewards, and safe hybrid-engine cache handling.

OraRL in One Update

Animated OraRL framework

An OraRL update separates reliable annotation guidance from on-policy normalization:

  1. Build the group: append one serialized annotation rollout to the policy samples generated for the same prompt.
  2. Keep the baseline on-policy: estimate the group baseline from policy rewards only.
  3. Guide and select: convert the annotation-policy reward gap into a correction, then retain a sign-balanced subset for the update.

This design uses task-native annotations directly and requires no chain-of-thought supervision or decoding.

Video-ORA Results

Video-ORA-9B results across seven task families

Dataset-Level Results

Dataset-level benchmark matrix comparing Video-ORA with multimodal baselines

Video-ORA-9B leads the matched seven-family comparison without CoT decoding. Best and second-best values are highlighted per row; † denotes an original-report value whose frame, prompt, split, or decoding settings may differ. Averages require complete family coverage.

Model Scaling

Animated Video-ORA model scaling from 0.8B to 9B

Data Scaling

Animated OraRL data scaling and reward dynamics

Models

ModelBackboneReleased recipeWeights
Video-ORA-9BQwen3.5-9Borarl_9b.yamlHugging Face
Video-ORA-4BQwen3.5-4Borarl_4b.yamlHugging Face

vLLM Serving

Both Video-ORA checkpoints load directly with vLLM 0.19.1 for OpenAI-compatible serving:

MODEL=OraRL/Video-ORA-9B

vllm serve "$MODEL" \
  --served-model-name Video-ORA-9B \
  --trust-remote-code \
  --dtype bfloat16 \
  --tensor-parallel-size 1 \
  --max-model-len 131072 \
  --limit-mm-per-prompt '{"image": 1, "video": 1}'

Set --tensor-parallel-size to the GPU count for multi-GPU deployment and lower --max-model-len on smaller-memory devices. Use enable_thinking=false in the chat template for answer-only inference.

Use OraRL

The release is organized around three user-facing workflows:

  1. Environment: install the pinned CUDA stack that covers both the bundled trainer and the evaluators.
  2. Training: prepare licensed local training data and launch GRPO or OraRL on one or multiple nodes.
  3. Evaluation: download Video-ORA and OraRL-Data, then run a smoke test or the complete paper suite.

Training and evaluation are dry runs by default; inspect the resolved command before adding --run. Checkpoints and evaluation media are hosted under the OraRL Hugging Face organization.

Acknowledgements

OraRL is built on veRL — a high-performance RL framework with HybridEngine. We thank its authors and contributors for open-sourcing the training infrastructure.

License

OraRL source is released under Apache-2.0. Datasets, models, benchmarks, and optional dependencies retain their original licenses; see NOTICE.

Citation

If you find OraRL useful, please consider giving this repository a ⭐ and citing our paper.

@article{li2026orarl,
  title   = {Annotations as Rollouts: Efficient and Scalable
             Reinforcement Learning for Video MLLMs},
  author  = {Li, Yunheng and Mu, Guohong and Li, Hao and
             Qian, Shengsheng and Zhang, Dingwen and Hou, Qibin
             and Cheng, Ming-Ming},
  journal = {arXiv preprint arXiv:2608.20492},
  year    = {2026},
  url     = {https://arxiv.org/abs/2608.20492}
}
multimodal-llm
reinforcement-learning
spatial-grounding
temporal-grounding
video-reasoning
video-segmentation
video-understanding
vision-language-model

HVision-NKU/OraRL

šŸŽ¬ OraRL — Annotations as Rollouts for efficient, scalable reinforcement learning of unified video MLLMs.

Python

163

6 commits

updated Aug 28, 2026

See the code

README

简体中文

OraRL

Annotations as Rollouts

Efficient and scalable reinforcement learning for unified video MLLMs

Yunheng Li Ā· Guohong Mu Ā· Hao Li Ā· Shengsheng Qian Ā· Dingwen Zhang Ā· Qibin Hou Ā· Ming-Ming Cheng

šŸ“„ Paper Ā Ā Ā·Ā Ā  🌐 Project Page Ā Ā Ā·Ā Ā  šŸŽ® Live Demo Ā Ā Ā·Ā Ā  šŸ¤— Models (4B / 9B)

āš™ļø Environment Ā Ā Ā·Ā Ā  šŸš€ Training Ā Ā Ā·Ā Ā  šŸ“Š Evaluation Ā Ā Ā·Ā Ā  āš–ļø License

Papers with Code: SOTA on ActivityNet-TimeLens Papers with Code: SOTA on Charades-TimeLens Papers with Code: SOTA on MeViS Papers with Code: SOTA on QVHighlights-TimeLens Validation Papers with Code: SOTA on VideoHolmes Papers with Code: #2 on ReasonVOS

View all 11 verified evaluations on Papers with Code

Animated OraRL method preview

ā–¶ Click the image to watch the 1:38 project overview.

Why OraRL

  • Annotation-as-rollout: annotations become reliable positive rollouts while policy samples retain an on-policy baseline.
  • Seven task families: one update rule covers temporal and spatial grounding, segmentation, tracking, spatial-temporal grounding, video QA, and spatial intelligence.
  • Efficient training (4B): sign-balanced pruning delivers 1.48Ɨ faster updates (92.5 → 62.4 s/step) while reducing peak per-GPU memory from 62.4 to 50.9 GB.
  • Efficient inference: on one H20 with vLLM in BF16, weight loading occupies 8.6 GiB (4B) and 17.6 GiB (9B). On ten-minute, 2-fps videos, answer-only decoding cuts median post-TTFT latency from 4.78 s to 130 ms and total latency from 29.03 to 24.30 s.
  • Multimodal veRL infrastructure: a unified video contract carries cached artifacts, raw paths, or inline frame tensors through vLLM rollouts and FSDP updates, with decode-once frame reuse, temporal metadata, task-grouped batching, asynchronous Ray rewards, and safe hybrid-engine cache handling.

OraRL in One Update

Animated OraRL framework

An OraRL update separates reliable annotation guidance from on-policy normalization:

  1. Build the group: append one serialized annotation rollout to the policy samples generated for the same prompt.
  2. Keep the baseline on-policy: estimate the group baseline from policy rewards only.
  3. Guide and select: convert the annotation-policy reward gap into a correction, then retain a sign-balanced subset for the update.

This design uses task-native annotations directly and requires no chain-of-thought supervision or decoding.

Video-ORA Results

Video-ORA-9B results across seven task families

Dataset-Level Results

Dataset-level benchmark matrix comparing Video-ORA with multimodal baselines

Video-ORA-9B leads the matched seven-family comparison without CoT decoding. Best and second-best values are highlighted per row; † denotes an original-report value whose frame, prompt, split, or decoding settings may differ. Averages require complete family coverage.

Model Scaling

Animated Video-ORA model scaling from 0.8B to 9B

Data Scaling

Animated OraRL data scaling and reward dynamics

Models

ModelBackboneReleased recipeWeights
Video-ORA-9BQwen3.5-9Borarl_9b.yamlHugging Face
Video-ORA-4BQwen3.5-4Borarl_4b.yamlHugging Face

vLLM Serving

Both Video-ORA checkpoints load directly with vLLM 0.19.1 for OpenAI-compatible serving:

MODEL=OraRL/Video-ORA-9B

vllm serve "$MODEL" \
  --served-model-name Video-ORA-9B \
  --trust-remote-code \
  --dtype bfloat16 \
  --tensor-parallel-size 1 \
  --max-model-len 131072 \
  --limit-mm-per-prompt '{"image": 1, "video": 1}'

Set --tensor-parallel-size to the GPU count for multi-GPU deployment and lower --max-model-len on smaller-memory devices. Use enable_thinking=false in the chat template for answer-only inference.

Use OraRL

The release is organized around three user-facing workflows:

  1. Environment: install the pinned CUDA stack that covers both the bundled trainer and the evaluators.
  2. Training: prepare licensed local training data and launch GRPO or OraRL on one or multiple nodes.
  3. Evaluation: download Video-ORA and OraRL-Data, then run a smoke test or the complete paper suite.

Training and evaluation are dry runs by default; inspect the resolved command before adding --run. Checkpoints and evaluation media are hosted under the OraRL Hugging Face organization.

Acknowledgements

OraRL is built on veRL — a high-performance RL framework with HybridEngine. We thank its authors and contributors for open-sourcing the training infrastructure.

License

OraRL source is released under Apache-2.0. Datasets, models, benchmarks, and optional dependencies retain their original licenses; see NOTICE.

Citation

If you find OraRL useful, please consider giving this repository a ⭐ and citing our paper.

@article{li2026orarl,
  title   = {Annotations as Rollouts: Efficient and Scalable
             Reinforcement Learning for Video MLLMs},
  author  = {Li, Yunheng and Mu, Guohong and Li, Hao and
             Qian, Shengsheng and Zhang, Dingwen and Hou, Qibin
             and Cheng, Ming-Ming},
  journal = {arXiv preprint arXiv:2608.20492},
  year    = {2026},
  url     = {https://arxiv.org/abs/2608.20492}
}
multimodal-llm
reinforcement-learning
spatial-grounding
temporal-grounding
video-reasoning
video-segmentation
video-understanding
vision-language-model

Languages

Python

92.9%

Shell

7.1%