InternRobotics/VLAC-Cut

Model

VLAC-Cut: Video Progress Estimation for Process-Level Robot Rollout Segmentation

1

67 commits

1 linked in READMEs

updated Jul 16, 2026

See the code

README

VLAC-Cut: Video Progress Estimation for Process-Level Robot Rollout Segmentation

Overview

VLAC-Cut is a process-level multimodal trajectory critic for robot post-training data curation. Given a natural-language task instruction, an optional task plan, and a robot rollout video, VLAC-Cut estimates signed task progress over time and identifies temporal segments associated with task advancement or degradation.

Unlike methods that assume task progress increases monotonically over time, VLAC-Cut models non-monotonic execution dynamics, including advancement, stagnation, regression, and recovery. This formulation supports process-level analysis of partial completion, temporary failure, subsequent recovery, and rollout segmentation for post-training data selection.

This Hugging Face repository contains the VLAC-Cut model weights and loading assets. The official inference examples and evaluation code are maintained in the GitHub repository.

Highlights

  • Video-level temporal reasoning: Analyzes robot execution videos rather than isolated images or image pairs and identifies temporal segments associated with task advancement or degradation.
  • Non-monotonic progress estimation: Captures advancement, stagnation, regression, and recovery without imposing a monotonically increasing progress assumption.
  • Zero-shot generalization: Generalizes across manipulation tasks, scenes, object configurations, and camera viewpoints.
  • Flexible temporal resolution: Supports configurable video sampling frequencies for both coarse- and fine-grained progress estimation.

Model Overview

PropertyDescription
Base modelQwen/Qwen3-VL-30B-A3B-Instruct
InputTask instruction, optional task plan, and sampled video frames
OutputTimestamped task-progress estimates
Sampling rate2 Hz-20 Hz
Default sampling rate2.0 Hz

Load with Transformers

from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "InternRobotics/VLAC-Cut"

processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    dtype="auto",
    device_map="auto",
)
model.eval()

Quick Start

Run progress inference on a local video using the GitHub code:

git clone https://github.com/InternRobotics/VLAC-cut
cd VLAC-cut

python scripts/run_example.py \
  --model-path InternRobotics/VLAC-Cut \
  --video-path <path-to-video.mp4> \
  --task-instruction "<natural-language task instruction>" \
  --task-plan $'<optional step-by-step task plan>' \
  --output-jsonl <path-to-output.jsonl>

Render a prediction JSONL file as an annotated video:

python scripts/utils/render_prediction_video.py \
  --input-jsonl <path-to-output.jsonl> \
  --output-video <path-to-preview.mp4>

Citation

Please cite the following paper when using VLAC-Cut, the released model, or the Video Progress Benchmark:

@misc{zhai2026helphumanefficientlargescalerobot,
      title={HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation}, 
      author={Shaopeng Zhai and Qi Zhang and Tianyi Zhang and Haoran Zhang and Fuxian Huang and Zhanhui Lin and Zijun Xu and Weinan Zhang},
      year={2026},
      eprint={2607.09776},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2607.09776}, 
}

License

The model weights and third-party training data may be subject to additional licenses or terms of use. The source code in the GitHub repository is released under the MIT License.

conversational
embodied-ai
endpoints_compatible
image-text-to-text
progress-estimation
qwen3-vl
qwen3_vl_moe
reward-modeling
robotics
safetensors
transformers
video-understanding

Contributors

futurefantasy

67 commits

InternRobotics/VLAC-Cut

Model

VLAC-Cut: Video Progress Estimation for Process-Level Robot Rollout Segmentation

1

67 commits

1 linked in READMEs

updated Jul 16, 2026

See the code

README

VLAC-Cut: Video Progress Estimation for Process-Level Robot Rollout Segmentation

Overview

VLAC-Cut is a process-level multimodal trajectory critic for robot post-training data curation. Given a natural-language task instruction, an optional task plan, and a robot rollout video, VLAC-Cut estimates signed task progress over time and identifies temporal segments associated with task advancement or degradation.

Unlike methods that assume task progress increases monotonically over time, VLAC-Cut models non-monotonic execution dynamics, including advancement, stagnation, regression, and recovery. This formulation supports process-level analysis of partial completion, temporary failure, subsequent recovery, and rollout segmentation for post-training data selection.

This Hugging Face repository contains the VLAC-Cut model weights and loading assets. The official inference examples and evaluation code are maintained in the GitHub repository.

Highlights

  • Video-level temporal reasoning: Analyzes robot execution videos rather than isolated images or image pairs and identifies temporal segments associated with task advancement or degradation.
  • Non-monotonic progress estimation: Captures advancement, stagnation, regression, and recovery without imposing a monotonically increasing progress assumption.
  • Zero-shot generalization: Generalizes across manipulation tasks, scenes, object configurations, and camera viewpoints.
  • Flexible temporal resolution: Supports configurable video sampling frequencies for both coarse- and fine-grained progress estimation.

Model Overview

PropertyDescription
Base modelQwen/Qwen3-VL-30B-A3B-Instruct
InputTask instruction, optional task plan, and sampled video frames
OutputTimestamped task-progress estimates
Sampling rate2 Hz-20 Hz
Default sampling rate2.0 Hz

Load with Transformers

from transformers import AutoModelForImageTextToText, AutoProcessor

model_id = "InternRobotics/VLAC-Cut"

processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
    model_id,
    dtype="auto",
    device_map="auto",
)
model.eval()

Quick Start

Run progress inference on a local video using the GitHub code:

git clone https://github.com/InternRobotics/VLAC-cut
cd VLAC-cut

python scripts/run_example.py \
  --model-path InternRobotics/VLAC-Cut \
  --video-path <path-to-video.mp4> \
  --task-instruction "<natural-language task instruction>" \
  --task-plan $'<optional step-by-step task plan>' \
  --output-jsonl <path-to-output.jsonl>

Render a prediction JSONL file as an annotated video:

python scripts/utils/render_prediction_video.py \
  --input-jsonl <path-to-output.jsonl> \
  --output-video <path-to-preview.mp4>

Citation

Please cite the following paper when using VLAC-Cut, the released model, or the Video Progress Benchmark:

@misc{zhai2026helphumanefficientlargescalerobot,
      title={HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation}, 
      author={Shaopeng Zhai and Qi Zhang and Tianyi Zhang and Haoran Zhang and Fuxian Huang and Zhanhui Lin and Zijun Xu and Weinan Zhang},
      year={2026},
      eprint={2607.09776},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2607.09776}, 
}

License

The model weights and third-party training data may be subject to additional licenses or terms of use. The source code in the GitHub repository is released under the MIT License.

conversational
embodied-ai
endpoints_compatible
image-text-to-text
progress-estimation
qwen3-vl
qwen3_vl_moe
reward-modeling
robotics
safetensors
transformers
video-understanding

Contributors

futurefantasy

67 commits