R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning.
Python
67
31 commits
updated May 14, 2025
Visual (Single) Object Tracking aims to continuously localize and estimate the scale of a target in subsequent video frames, given only its initial state in the first frame. This task can be simplified to template matching between image pairs, with traditional trackers predominantly employing explicit classification-regression modeling through Correlation Filters, Siamese networks, and Vision Transformers (ViT). Leveraging advancements in Multi-Modal Large Language Models (MLLMs) such as Qwen2.5-VL and their robust grounding capabilities, we explore adopting MLLMs for end-to-end tracking tasks, eliminating the need for fragmented subtask modeling.
R1-Track supports flexible initialization from either text descriptions or bounding boxes.
{
"role": "user",
"content": [
{
"type": "image",
"image": xxx,
},
{
"type": "image",
"image": xxx,
},
{"type": "text", "text": "Please identify the target specified by the bounding box [241,66,329,154] in the first image and locate it in the second image. \n Return the coordinates in [x_min,y_min,x_max,y_max] format."},
# R1-Track-100k:
#Given two images, you need to:\n1. Analyze and Identify the target object marked by bounding box <BBOXFLAG> in <image_1>;\n2. Re-locate this target in <image_2>;\n3. Return [x_min, y_min, x_max, y_max] coordinates of the target in <image_2>.
],
}
{
"role": "user",
"content": [
{
"type": "image",
"image": xxx,
},
{
"type": "image",
"image": xxx,
},
{"type": "text", "text": "You FIRST think about the reasoning process as an internal monologue and then provide the final answer. \n The reasoning process MUST BE enclosed within <think> </think> tags. The final answer MUST BE put in <answer> </answer> tags. \n Please identify the target specified by the bounding box [241,66,329,154] in the first image and locate it in the second image. \n Return the coordinates in [x_min,y_min,x_max,y_max] format."},
# R1-Track-100k:
#You FIRST think about the reasoning process as an internal monologue and then provide the final answer. \n The reasoning process MUST BE enclosed within <think> </think> tags. The final answer MUST BE put in <answer> </answer> tags. \n Given two images, you need to:\n1. Analyze and Identify the target object marked by bounding box <BBOXFLAG> in <image_1>;\n2. Re-locate this target in <image_2>;\n3. Return [x_min, y_min, x_max, y_max] coordinates of the target in <image_2>.
],
}
python generation_data/sampling_data.py
python generation_data/gen_huggingface_data.py
Please refer to the official LLaMA-Factory repo for env configuration guidelines.
CUDA_VISIBLE_DEVICES=0,1,2,3 llamafactory-cli train examples/train_lora/r1_track_lora_sft.yaml
llamafactory-cli export examples/merge_lora/r1_track_lora_sft.yaml
Please refer to the official EasyR1 repo for env configuration guidelines.
cd EasyR1
bash examples/qwen2_5_vl_3b_track5k_grpo_w_think.sh
CUDA_VISIBLE_DEVICES=0,1,2,3 python -m vllm.entrypoints.openai.api_server --served-model-name R1-Track --model WangBiao/R1-Track-GRPO --gpu-memory-utilization 0.9 --tensor-parallel-size 4 --port 8888 --limit-mm-per-prompt image=2
python infer_script/r1track.py
| Tracker/GOT10k | Crop Size | Finetune Data | $AO$ | $SR_{0.5}$ | $SR_{0.75}$ | Params | Ckpt |
|---|---|---|---|---|---|---|---|
| LLaVA-1.5 | - | - | 0.235 | 0.202 | - | 7B | - |
| Qwen2.5-VL-7B-Instruct | - | - | 0.126 | 0.011 | - | 7B | - |
| Qwen2.5-VL-72B-Instruct | - | - | - | - | - | 72B | - |
| R1-Track-SFT | $336 \times 336$ | R1-Track-5K | 0.543 | 0.633 | 0.338 | 3B | R1-Track-SFT |
| R1-Track-GRPO | $336 \times 336$ | R1-Track-5K | 0.586 | 0.676 | 0.470 | 3B | R1-Track-GRPO |
| R1-Track-GRPO-wo-Think | $336 \times 336$ | R1-Track-5k | 0.585 | 0.673 | 0.500 | 3B | R1-Track-GRPO-wo-Think |
| R1-Track-SFT | $336 \times 336$ | R1-Track-100k | 0.667 | 0.746 | 0.620 | 3B | R1-Track-SFT-0503-lora |
| R1-Track-GRPO | $336 \times 336$ | R1-Track-100k | 0.672 | 0.759 | 0.624 | 3B | - |
| R1-Track-GRPO-wo-Think | $336 \times 336$ | R1-Track-100k | 0.680 | 0.766 | 0.637 | 3B | R1-Track-GRPO-wo-Think-0503 |
Note: In our experiment, we found that letting the 3b base model (without cold-start by COT data) directly output the result instead of following <think></think><answer></answer> would lead to a much higher score on GOT-10k.
We will strive to elevate R1-Track to the T0 level of trackers.
Please open an issue if you have any questions.
@misc{wang2025r1track,
title = {R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning},
author = {Biao Wang},
howpublished = {\url{https://github.com/Wangbiao2/R1-Track}},
year = {2025}
}
R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning.
Python
67
31 commits
updated May 14, 2025
Visual (Single) Object Tracking aims to continuously localize and estimate the scale of a target in subsequent video frames, given only its initial state in the first frame. This task can be simplified to template matching between image pairs, with traditional trackers predominantly employing explicit classification-regression modeling through Correlation Filters, Siamese networks, and Vision Transformers (ViT). Leveraging advancements in Multi-Modal Large Language Models (MLLMs) such as Qwen2.5-VL and their robust grounding capabilities, we explore adopting MLLMs for end-to-end tracking tasks, eliminating the need for fragmented subtask modeling.
R1-Track supports flexible initialization from either text descriptions or bounding boxes.
{
"role": "user",
"content": [
{
"type": "image",
"image": xxx,
},
{
"type": "image",
"image": xxx,
},
{"type": "text", "text": "Please identify the target specified by the bounding box [241,66,329,154] in the first image and locate it in the second image. \n Return the coordinates in [x_min,y_min,x_max,y_max] format."},
# R1-Track-100k:
#Given two images, you need to:\n1. Analyze and Identify the target object marked by bounding box <BBOXFLAG> in <image_1>;\n2. Re-locate this target in <image_2>;\n3. Return [x_min, y_min, x_max, y_max] coordinates of the target in <image_2>.
],
}
{
"role": "user",
"content": [
{
"type": "image",
"image": xxx,
},
{
"type": "image",
"image": xxx,
},
{"type": "text", "text": "You FIRST think about the reasoning process as an internal monologue and then provide the final answer. \n The reasoning process MUST BE enclosed within <think> </think> tags. The final answer MUST BE put in <answer> </answer> tags. \n Please identify the target specified by the bounding box [241,66,329,154] in the first image and locate it in the second image. \n Return the coordinates in [x_min,y_min,x_max,y_max] format."},
# R1-Track-100k:
#You FIRST think about the reasoning process as an internal monologue and then provide the final answer. \n The reasoning process MUST BE enclosed within <think> </think> tags. The final answer MUST BE put in <answer> </answer> tags. \n Given two images, you need to:\n1. Analyze and Identify the target object marked by bounding box <BBOXFLAG> in <image_1>;\n2. Re-locate this target in <image_2>;\n3. Return [x_min, y_min, x_max, y_max] coordinates of the target in <image_2>.
],
}
python generation_data/sampling_data.py
python generation_data/gen_huggingface_data.py
Please refer to the official LLaMA-Factory repo for env configuration guidelines.
CUDA_VISIBLE_DEVICES=0,1,2,3 llamafactory-cli train examples/train_lora/r1_track_lora_sft.yaml
llamafactory-cli export examples/merge_lora/r1_track_lora_sft.yaml
Please refer to the official EasyR1 repo for env configuration guidelines.
cd EasyR1
bash examples/qwen2_5_vl_3b_track5k_grpo_w_think.sh
CUDA_VISIBLE_DEVICES=0,1,2,3 python -m vllm.entrypoints.openai.api_server --served-model-name R1-Track --model WangBiao/R1-Track-GRPO --gpu-memory-utilization 0.9 --tensor-parallel-size 4 --port 8888 --limit-mm-per-prompt image=2
python infer_script/r1track.py
| Tracker/GOT10k | Crop Size | Finetune Data | $AO$ | $SR_{0.5}$ | $SR_{0.75}$ | Params | Ckpt |
|---|---|---|---|---|---|---|---|
| LLaVA-1.5 | - | - | 0.235 | 0.202 | - | 7B | - |
| Qwen2.5-VL-7B-Instruct | - | - | 0.126 | 0.011 | - | 7B | - |
| Qwen2.5-VL-72B-Instruct | - | - | - | - | - | 72B | - |
| R1-Track-SFT | $336 \times 336$ | R1-Track-5K | 0.543 | 0.633 | 0.338 | 3B | R1-Track-SFT |
| R1-Track-GRPO | $336 \times 336$ | R1-Track-5K | 0.586 | 0.676 | 0.470 | 3B | R1-Track-GRPO |
| R1-Track-GRPO-wo-Think | $336 \times 336$ | R1-Track-5k | 0.585 | 0.673 | 0.500 | 3B | R1-Track-GRPO-wo-Think |
| R1-Track-SFT | $336 \times 336$ | R1-Track-100k | 0.667 | 0.746 | 0.620 | 3B | R1-Track-SFT-0503-lora |
| R1-Track-GRPO | $336 \times 336$ | R1-Track-100k | 0.672 | 0.759 | 0.624 | 3B | - |
| R1-Track-GRPO-wo-Think | $336 \times 336$ | R1-Track-100k | 0.680 | 0.766 | 0.637 | 3B | R1-Track-GRPO-wo-Think-0503 |
Note: In our experiment, we found that letting the 3b base model (without cold-start by COT data) directly output the result instead of following <think></think><answer></answer> would lead to a much higher score on GOT-10k.
We will strive to elevate R1-Track to the T0 level of trackers.
Please open an issue if you have any questions.
@misc{wang2025r1track,
title = {R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning},
author = {Biao Wang},
howpublished = {\url{https://github.com/Wangbiao2/R1-Track}},
year = {2025}
}