VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
1
4 commits
2 linked in READMEs
updated Dec 31, 2024
VideoRefer-Bench is a comprehensive benchmark to evaluate the object-level video understanding capabilities of a model, which consists of two sub-benchmarks: VideoRefer-Bench-D and VideoRefer-Bench-Q.
The benchmark is designed to evaluate the description generation performance of video-based referring models. The benchmark comprises a total of 400 curated data entries. We curated the test set based on Panda-70M, employing the automatic pipeline, followed by a meticulous human check.
This benchmark covers four key aspects:
| Type | GPT-4o | InternVL2-26B | Qwen2-VL-7B | Elysium | Artemis | VideoRefer |
|---|---|---|---|---|---|---|
| Subject Correspondence | 3.34/4.15 | 3.55/4.08 | 2.97/3.30 | 2.35/- | -/3.42 | 4.41/4.44 |
| Appearance Description | 2.96/3.31 | 2.99/3.35 | 2.24/2.54 | 0.30/- | -/1.34 | 3.27/3.27 |
| Temporal Description | 3.01/3.11 | 2.57/3.08 | 2.03/2.22 | 0.02/- | -/1.39 | 3.03/3.10 |
| Hallucinaton Detection | 2.50/2.43 | 2.25/2.28 | 2.31/2.12 | 3.59/- | -/2.90 | 2.97/3.04 |
| Average | 2.95/3.25 | 2.84/3.20 | 2.39/2.55 | 1.57/- | -/2.26 | 3.42/3.46 |
For each object, we uniformly sampled 32 frames to generate the corresponding mask.
The data format organized in the benchmark json file is as below:
[
{
"id": 0,
"video": "rLlzmcp3J6s_0:01:09.633_0:01:14.333.mp4",
"caption": "The cub is a smaller, light colored lion. It is lying down and resting its head against the other lion. The cub looks calm and relaxed. It is the lion on the far left side of the frame.",
"frame_idx": "36",
"annotation":[
{
"2":{
"segmentation": {
}
},
"6":{
"segmentation": {
}
},
...
}
]
}
]
frame_idx: When using single-frame mask mode, we only use the single mask with the frame_idx.RLE format.The benchmark is designed to evaluate the proficiency of MLLMs in interpreting video objects, including 1,000 high-quality multiple-choice questions. The benchmark covers five types of questions:
For each object, we uniformly sampled 32 frames to generate the corresponding mask. The data format organized in the benchmark json file is as below:
[
{
"id": 0,
"video": "DAVIS/JPEGImages/480p/aerobatics",
"Question": "What is <object3><region> not wearing?",
"type": "Basic Questions",
"options": [
"(A) A helmet",
"(B) A hat",
"(C) Sunglasses",
"(D) A watch"
],
"Answer": "(A) A helmet",
"frame_idx": "57",
"annotation":[
{
"0":{
"segmentation": {
}
},
"3":{
"segmentation": {
}
},
...
}
]
}
]
frame_idx: When using single-frame mask mode, we only use the single mask with the frame_idx.RLE format.VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
1
4 commits
2 linked in READMEs
updated Dec 31, 2024
VideoRefer-Bench is a comprehensive benchmark to evaluate the object-level video understanding capabilities of a model, which consists of two sub-benchmarks: VideoRefer-Bench-D and VideoRefer-Bench-Q.
The benchmark is designed to evaluate the description generation performance of video-based referring models. The benchmark comprises a total of 400 curated data entries. We curated the test set based on Panda-70M, employing the automatic pipeline, followed by a meticulous human check.
This benchmark covers four key aspects:
| Type | GPT-4o | InternVL2-26B | Qwen2-VL-7B | Elysium | Artemis | VideoRefer |
|---|---|---|---|---|---|---|
| Subject Correspondence | 3.34/4.15 | 3.55/4.08 | 2.97/3.30 | 2.35/- | -/3.42 | 4.41/4.44 |
| Appearance Description | 2.96/3.31 | 2.99/3.35 | 2.24/2.54 | 0.30/- | -/1.34 | 3.27/3.27 |
| Temporal Description | 3.01/3.11 | 2.57/3.08 | 2.03/2.22 | 0.02/- | -/1.39 | 3.03/3.10 |
| Hallucinaton Detection | 2.50/2.43 | 2.25/2.28 | 2.31/2.12 | 3.59/- | -/2.90 | 2.97/3.04 |
| Average | 2.95/3.25 | 2.84/3.20 | 2.39/2.55 | 1.57/- | -/2.26 | 3.42/3.46 |
For each object, we uniformly sampled 32 frames to generate the corresponding mask.
The data format organized in the benchmark json file is as below:
[
{
"id": 0,
"video": "rLlzmcp3J6s_0:01:09.633_0:01:14.333.mp4",
"caption": "The cub is a smaller, light colored lion. It is lying down and resting its head against the other lion. The cub looks calm and relaxed. It is the lion on the far left side of the frame.",
"frame_idx": "36",
"annotation":[
{
"2":{
"segmentation": {
}
},
"6":{
"segmentation": {
}
},
...
}
]
}
]
frame_idx: When using single-frame mask mode, we only use the single mask with the frame_idx.RLE format.The benchmark is designed to evaluate the proficiency of MLLMs in interpreting video objects, including 1,000 high-quality multiple-choice questions. The benchmark covers five types of questions:
For each object, we uniformly sampled 32 frames to generate the corresponding mask. The data format organized in the benchmark json file is as below:
[
{
"id": 0,
"video": "DAVIS/JPEGImages/480p/aerobatics",
"Question": "What is <object3><region> not wearing?",
"type": "Basic Questions",
"options": [
"(A) A helmet",
"(B) A hat",
"(C) Sunglasses",
"(D) A watch"
],
"Answer": "(A) A helmet",
"frame_idx": "57",
"annotation":[
{
"0":{
"segmentation": {
}
},
"3":{
"segmentation": {
}
},
...
}
]
}
]
frame_idx: When using single-frame mask mode, we only use the single mask with the frame_idx.RLE format.