Built by Rapidata.
This dataset contains 281,738 human responses, collected with the Rapidata Python SDK, comparing how well 14 image-to-video models and world models execute a described camera movement from a single still image. Each row is a head-to-head comparison between two models' clips generated from the same still and the same instruction, judged by human annotators who watched a reference animation of the requested movement.
The task is deliberately narrow: the scene must stay exactly as it is, and only the camera may move β a dolly, a truck, a pedestal, a pan, a tilt, a roll, a zoom, or a combination of them. Every clip is ranked purely by human judgement, no automated metrics.
If you get value from this dataset and would like to see more in the future, please consider liking it β€οΈ
To evaluate your own models and create a leaderboard, check out our MRI.
Explore the full interactive leaderboard β filter by model, movement, scene and number of motions, and inspect individual head-to-head matchups on Rapidata.
π Click to open and interact with it on rapidata.ai
| Head-to-head comparisons (rows) | 42,481 |
| Human votes | 281,738 (at least 5 per pair, 6.6 on average) |
| Models compared | 14 (8 commercial image-to-video models Β· 6 world models) |
| Camera movements | 30 (14 single moves Β· 13 two-move combinations Β· 3 three-move chains) |
| Input stills | 22 |
| Prompts (still Γ movement, hand-picked) | 600 |
| Unique generated clips | 8,400 |
Every model is compared on the same kind of clip pairs on one question, run as one leaderboard on a Rapidata benchmark:
| Leaderboard | Question shown to annotators | Prompt shown? | Measures |
|---|---|---|---|
| Movement | "Which video better matches the movement of the top video?" | No β annotators see the reference clip of the movement instead of the sentence | whether the camera did what was asked, and nothing else |
Annotators are shown the movement's reference animation on top and the two generated clips below it. They never read the text instruction: the prompts are written for a generator (degrees, constraints, frames of reference), while the reference clip shows a human unambiguously what the camera is supposed to do.
.vertical-container { display: flex; flex-direction: column; gap: 60px; } .horizontal-container { display: flex; flex-direction: row; justify-content: center; gap: 60px; } .media-container video, .media-container img { max-height: 260px; margin: 0; object-fit: contain; width: auto; box-sizing: content-box; } .media-container { display: flex; justify-content: space-around; align-items: flex-start; gap: .5rem } .container { width: 90%; margin: 0 auto; } .text-center { text-align: center; } .score-amount { margin: 0; margin-top: 10px; } .score-percentage { font-size: 12px; font-weight: semi-bold; }Each example shows the input still, the reference clip annotators were shown, and the two generated clips with the share of the (userScore-weighted) votes each one received. The green border marks the winner.
The camera tilts thirty degrees downward, staying in place.
The camera rises while turning thirty degrees to the right.
The camera rises, then tilts twenty degrees downward, then moves forward at the same height.
| # | Model | Lab | Type | ELO |
|---|---|---|---|---|
| 1 | Gemini Omni Flash 1.1 | commercial image-to-video | 1284.2 | |
| 2 | Wan 3.0 | Alibaba Cloud | commercial image-to-video | 1249.2 |
| 3 | Happy Horse 1.1 | Alibaba | commercial image-to-video | 1181.1 |
| 4 | Sana WM | NVIDIA | world model | 1137.0 |
| 5 | Alaya World | Alaya Lab | world model | 1118.1 |
| 6 | Kling V3 Pro | Kuaishou | commercial image-to-video | 1062.2 |
| 7 | Grok Imagine Video 1.5 | xAI | commercial image-to-video | 1039.3 |
| 8 | Dream X World | AMAP | world model | 1003.6 |
| 9 | Cosmos 3 Nano | NVIDIA | world model | 973.8 |
| 10 | MiniMax H3 | MiniMax | commercial image-to-video | 954.0 |
| 11 | Veo 3.1 | commercial image-to-video | 917.9 | |
| 12 | Vidu Q3 | Shengshu AI | commercial image-to-video | 877.6 |
| 13 | Yume 1.5 | Shanghai AI Laboratory | world model | 812.7 |
| 14 | Cosmos Predict 2.5 | NVIDIA | world model | 770.2 |
The live leaderboard additionally lets you slice the standings by number of motions (1 / 2 / 3), by the individual motion a prompt contains (move forward, tilt down, roll clockwise, β¦) and by scene, so you can see which models hold up as the instruction gets harder and which axes they drop.
This section is intentionally detailed so the generation process is fully reproducible and transparent.
The movements cover all six rigid-body degrees of freedom of a camera, each in both directions, plus zoom β and then combine them. They are graded into three difficulty tiers, which is the axis the benchmark reports on:
| Tier | Movements | Prompts | What it tests |
|---|---|---|---|
| 1 motion | 14 | 289 | one named move: dolly in / out, truck left / right, pedestal up / down, pan left 90Β° / right 45Β°, tilt up / down 30Β°, roll clockwise / counter-clockwise 45Β°, zoom in / out |
| 2 motions | 13 | 251 | two moves β either simultaneous ("rises while turning thirty degrees to the right": control decoupling) or sequential ("slides to the left, then turns thirty degrees to the right": planning) |
| 3 motions | 3 | 60 | three moves coupled or chained, e.g. "rises, then tilts twenty degrees downward, then moves forward at the same height" |
The wording follows strict rules so that every sentence isolates one skill:
while means simultaneous and then means sequential, and neither word is used for anything else.
The two-move tier is split between the two, because they fail independently.turns (yaw) and rolls (lens axis) deliberately keep
different verbs.Every two- and three-move prompt is composed of one-move prompts from the same set, so a failure can be pinned to a dropped axis: does a model that passes every part still fail the chain?
The stills are copyright-free photographs (one CG interior render) from Pexels and Pixabay, chosen for built environments people inhabit β a mirrored room, a Gothic nave, a library hall, a supermarket aisle, a crowded public square, a car park, a bowling alley, a workshop, a conservatory, a hotel corridor β plus a handful of landscapes and landmarks (Lake Louise, the Canadian Rockies, the Sphinx, the Colosseum, Manhattan from above). Each was picked for a named hard property: parallax that cannot be faked by cropping, populated vertical structure so a rise is measurable, repeated bays or shelves that act as a ruler, mirrors that contradict a wrong reflection.
Not every movement suits every still, so the prompt-to-image mapping is hand-picked rather than a cross product: a descend needs headroom below the lens, a wide pan needs something worth looking at off-axis, a tilt down on a near-nadir aerial reaches straight-down where the scene ends. The result is 600 prompts β between 17 and 30 movements per still.
A sentence like "the camera rolls forty-five degrees clockwise, staying in place" is precise but not
self-evident, so each movement gets a short looping GIF that shows an annotator the movement before they judge
whether a model performed it. The clips are rendered from scratch with a small numpy raytracer, and all 30 share
one synthetic scene, so the only thing that differs between two clips is the camera. Rotations, zooms and
compound moves visibly decelerate and stop at their stated angle; sequential (then) moves get a beat of
stillness at the seam, which is the only thing separating "one, then the other" from "both at once" in a picture.
The reference_gif column links this clip for every row.
Each still and its selected instructions were sent to every model as an image-to-video request. Two very different kinds of model compete on the same footing:
Eight commercial image-to-video models, generated at 1080p / 5 s through their public API endpoints with default settings β prompt expansion included, because the point is to compare each model as it comes, not as we tuned it. Audio was switched off wherever there was a switch, since nothing is graded on sound.
Google β Gemini Omni Flash 1.1, Veo 3.1 Β· Alibaba β Wan 3.0, Happy Horse 1.1 Β· Kuaishou β Kling V3 Pro Β· MiniMax β H3 Β· xAI β Grok Imagine Video 1.5 Β· Shengshu AI β Vidu Q3
Six research world models, run on their own GPUs with camera-controlled builds where the model exposes camera conditioning, at the models' native output (81 frames at 16 fps for most, 128 frames at 24 fps for Alaya World, 121 frames at 24 fps for Cosmos 3 Nano β all about 5 s). Every model was handed the same seed for the same prompt, derived deterministically from the prompt identifier.
NVIDIA β Cosmos 3 Nano, Cosmos Predict 2.5, Sana WM Β· Alaya Lab β Alaya World Β· AMAP β Dream X World Β· Shanghai AI Laboratory β Yume 1.5
Every model produced a clip for all 600 prompts, so each model meets every other one on identical inputs.
The clips were uploaded to a Rapidata MRI benchmark as one participant per model,
and the Movement leaderboard was run over the resulting pairwise matchups. Annotators came from a qualified
audience: before seeing any benchmark pair, each one had to pass a gate of known-answer clip pairs asking the
same question, at 75% accuracy or better. Every pair received at least 5 votes (6.6 on average), and per-annotator
detail β chosen side, country, language, gender, age bucket, occupation and the annotator's userScore β is
preserved in the detailed_results_movement column. Votes came from annotators in 128 countries.
One row per head-to-head clip pair generated from the same still and the same instruction.
| Column | Type | Description |
|---|---|---|
prompt | string | the camera-movement instruction both clips were generated from |
reference_gif | string | public URL of the reference clip annotators were shown instead of the text |
input_image | string | public URL of the still both clips were generated from |
video1 | string | public URL of the clip from model1 |
video2 | string | public URL of the clip from model2 |
model1 | string | name of the model that produced video1 |
model2 | string | name of the model that produced video2 |
weighted_results_video1_movement | float32 | sum of the userScore weights of the votes for video1 |
weighted_results_video2_movement | float32 | sum of the userScore weights of the votes for video2 |
detailed_results_movement | string | every vote on the pair as a JSON array (votedFor is A for video1 and B for video2, plus annotator demographics, userScore and votedAt) |
The two weighted_results_* values are userScore-weighted vote sums, not probabilities β they do not sum to 1.
Divide by their sum for a normalised win share, which is what the example percentages above show.
import json
from datasets import load_dataset
ds = load_dataset("Rapidata/camera-movement", split="train")
row = ds[0]
print(row["prompt"], "|", row["model1"], "vs", row["model2"])
w1, w2 = row["weighted_results_video1_movement"], row["weighted_results_video2_movement"]
print(f"win share {row['model1']}: {w1 / (w1 + w2):.0%}")
votes = json.loads(row["detailed_results_movement"])
print(len(votes), "votes, first one:", votes[0])
This dataset combines material under different terms:
Rapidata's technology makes collecting human feedback at scale faster and more accessible than ever before. Visit rapidata.ai to learn more about how we're revolutionizing human feedback collection for AI development.
Explore our latest model rankings on our website.
Built by Rapidata.
This dataset contains 281,738 human responses, collected with the Rapidata Python SDK, comparing how well 14 image-to-video models and world models execute a described camera movement from a single still image. Each row is a head-to-head comparison between two models' clips generated from the same still and the same instruction, judged by human annotators who watched a reference animation of the requested movement.
The task is deliberately narrow: the scene must stay exactly as it is, and only the camera may move β a dolly, a truck, a pedestal, a pan, a tilt, a roll, a zoom, or a combination of them. Every clip is ranked purely by human judgement, no automated metrics.
If you get value from this dataset and would like to see more in the future, please consider liking it β€οΈ
To evaluate your own models and create a leaderboard, check out our MRI.
Explore the full interactive leaderboard β filter by model, movement, scene and number of motions, and inspect individual head-to-head matchups on Rapidata.
π Click to open and interact with it on rapidata.ai
| Head-to-head comparisons (rows) | 42,481 |
| Human votes | 281,738 (at least 5 per pair, 6.6 on average) |
| Models compared | 14 (8 commercial image-to-video models Β· 6 world models) |
| Camera movements | 30 (14 single moves Β· 13 two-move combinations Β· 3 three-move chains) |
| Input stills | 22 |
| Prompts (still Γ movement, hand-picked) | 600 |
| Unique generated clips | 8,400 |
Every model is compared on the same kind of clip pairs on one question, run as one leaderboard on a Rapidata benchmark:
| Leaderboard | Question shown to annotators | Prompt shown? | Measures |
|---|---|---|---|
| Movement | "Which video better matches the movement of the top video?" | No β annotators see the reference clip of the movement instead of the sentence | whether the camera did what was asked, and nothing else |
Annotators are shown the movement's reference animation on top and the two generated clips below it. They never read the text instruction: the prompts are written for a generator (degrees, constraints, frames of reference), while the reference clip shows a human unambiguously what the camera is supposed to do.
.vertical-container { display: flex; flex-direction: column; gap: 60px; } .horizontal-container { display: flex; flex-direction: row; justify-content: center; gap: 60px; } .media-container video, .media-container img { max-height: 260px; margin: 0; object-fit: contain; width: auto; box-sizing: content-box; } .media-container { display: flex; justify-content: space-around; align-items: flex-start; gap: .5rem } .container { width: 90%; margin: 0 auto; } .text-center { text-align: center; } .score-amount { margin: 0; margin-top: 10px; } .score-percentage { font-size: 12px; font-weight: semi-bold; }Each example shows the input still, the reference clip annotators were shown, and the two generated clips with the share of the (userScore-weighted) votes each one received. The green border marks the winner.
The camera tilts thirty degrees downward, staying in place.
The camera rises while turning thirty degrees to the right.
The camera rises, then tilts twenty degrees downward, then moves forward at the same height.
| # | Model | Lab | Type | ELO |
|---|---|---|---|---|
| 1 | Gemini Omni Flash 1.1 | commercial image-to-video | 1284.2 | |
| 2 | Wan 3.0 | Alibaba Cloud | commercial image-to-video | 1249.2 |
| 3 | Happy Horse 1.1 | Alibaba | commercial image-to-video | 1181.1 |
| 4 | Sana WM | NVIDIA | world model | 1137.0 |
| 5 | Alaya World | Alaya Lab | world model | 1118.1 |
| 6 | Kling V3 Pro | Kuaishou | commercial image-to-video | 1062.2 |
| 7 | Grok Imagine Video 1.5 | xAI | commercial image-to-video | 1039.3 |
| 8 | Dream X World | AMAP | world model | 1003.6 |
| 9 | Cosmos 3 Nano | NVIDIA | world model | 973.8 |
| 10 | MiniMax H3 | MiniMax | commercial image-to-video | 954.0 |
| 11 | Veo 3.1 | commercial image-to-video | 917.9 | |
| 12 | Vidu Q3 | Shengshu AI | commercial image-to-video | 877.6 |
| 13 | Yume 1.5 | Shanghai AI Laboratory | world model | 812.7 |
| 14 | Cosmos Predict 2.5 | NVIDIA | world model | 770.2 |
The live leaderboard additionally lets you slice the standings by number of motions (1 / 2 / 3), by the individual motion a prompt contains (move forward, tilt down, roll clockwise, β¦) and by scene, so you can see which models hold up as the instruction gets harder and which axes they drop.
This section is intentionally detailed so the generation process is fully reproducible and transparent.
The movements cover all six rigid-body degrees of freedom of a camera, each in both directions, plus zoom β and then combine them. They are graded into three difficulty tiers, which is the axis the benchmark reports on:
| Tier | Movements | Prompts | What it tests |
|---|---|---|---|
| 1 motion | 14 | 289 | one named move: dolly in / out, truck left / right, pedestal up / down, pan left 90Β° / right 45Β°, tilt up / down 30Β°, roll clockwise / counter-clockwise 45Β°, zoom in / out |
| 2 motions | 13 | 251 | two moves β either simultaneous ("rises while turning thirty degrees to the right": control decoupling) or sequential ("slides to the left, then turns thirty degrees to the right": planning) |
| 3 motions | 3 | 60 | three moves coupled or chained, e.g. "rises, then tilts twenty degrees downward, then moves forward at the same height" |
The wording follows strict rules so that every sentence isolates one skill:
while means simultaneous and then means sequential, and neither word is used for anything else.
The two-move tier is split between the two, because they fail independently.turns (yaw) and rolls (lens axis) deliberately keep
different verbs.Every two- and three-move prompt is composed of one-move prompts from the same set, so a failure can be pinned to a dropped axis: does a model that passes every part still fail the chain?
The stills are copyright-free photographs (one CG interior render) from Pexels and Pixabay, chosen for built environments people inhabit β a mirrored room, a Gothic nave, a library hall, a supermarket aisle, a crowded public square, a car park, a bowling alley, a workshop, a conservatory, a hotel corridor β plus a handful of landscapes and landmarks (Lake Louise, the Canadian Rockies, the Sphinx, the Colosseum, Manhattan from above). Each was picked for a named hard property: parallax that cannot be faked by cropping, populated vertical structure so a rise is measurable, repeated bays or shelves that act as a ruler, mirrors that contradict a wrong reflection.
Not every movement suits every still, so the prompt-to-image mapping is hand-picked rather than a cross product: a descend needs headroom below the lens, a wide pan needs something worth looking at off-axis, a tilt down on a near-nadir aerial reaches straight-down where the scene ends. The result is 600 prompts β between 17 and 30 movements per still.
A sentence like "the camera rolls forty-five degrees clockwise, staying in place" is precise but not
self-evident, so each movement gets a short looping GIF that shows an annotator the movement before they judge
whether a model performed it. The clips are rendered from scratch with a small numpy raytracer, and all 30 share
one synthetic scene, so the only thing that differs between two clips is the camera. Rotations, zooms and
compound moves visibly decelerate and stop at their stated angle; sequential (then) moves get a beat of
stillness at the seam, which is the only thing separating "one, then the other" from "both at once" in a picture.
The reference_gif column links this clip for every row.
Each still and its selected instructions were sent to every model as an image-to-video request. Two very different kinds of model compete on the same footing:
Eight commercial image-to-video models, generated at 1080p / 5 s through their public API endpoints with default settings β prompt expansion included, because the point is to compare each model as it comes, not as we tuned it. Audio was switched off wherever there was a switch, since nothing is graded on sound.
Google β Gemini Omni Flash 1.1, Veo 3.1 Β· Alibaba β Wan 3.0, Happy Horse 1.1 Β· Kuaishou β Kling V3 Pro Β· MiniMax β H3 Β· xAI β Grok Imagine Video 1.5 Β· Shengshu AI β Vidu Q3
Six research world models, run on their own GPUs with camera-controlled builds where the model exposes camera conditioning, at the models' native output (81 frames at 16 fps for most, 128 frames at 24 fps for Alaya World, 121 frames at 24 fps for Cosmos 3 Nano β all about 5 s). Every model was handed the same seed for the same prompt, derived deterministically from the prompt identifier.
NVIDIA β Cosmos 3 Nano, Cosmos Predict 2.5, Sana WM Β· Alaya Lab β Alaya World Β· AMAP β Dream X World Β· Shanghai AI Laboratory β Yume 1.5
Every model produced a clip for all 600 prompts, so each model meets every other one on identical inputs.
The clips were uploaded to a Rapidata MRI benchmark as one participant per model,
and the Movement leaderboard was run over the resulting pairwise matchups. Annotators came from a qualified
audience: before seeing any benchmark pair, each one had to pass a gate of known-answer clip pairs asking the
same question, at 75% accuracy or better. Every pair received at least 5 votes (6.6 on average), and per-annotator
detail β chosen side, country, language, gender, age bucket, occupation and the annotator's userScore β is
preserved in the detailed_results_movement column. Votes came from annotators in 128 countries.
One row per head-to-head clip pair generated from the same still and the same instruction.
| Column | Type | Description |
|---|---|---|
prompt | string | the camera-movement instruction both clips were generated from |
reference_gif | string | public URL of the reference clip annotators were shown instead of the text |
input_image | string | public URL of the still both clips were generated from |
video1 | string | public URL of the clip from model1 |
video2 | string | public URL of the clip from model2 |
model1 | string | name of the model that produced video1 |
model2 | string | name of the model that produced video2 |
weighted_results_video1_movement | float32 | sum of the userScore weights of the votes for video1 |
weighted_results_video2_movement | float32 | sum of the userScore weights of the votes for video2 |
detailed_results_movement | string | every vote on the pair as a JSON array (votedFor is A for video1 and B for video2, plus annotator demographics, userScore and votedAt) |
The two weighted_results_* values are userScore-weighted vote sums, not probabilities β they do not sum to 1.
Divide by their sum for a normalised win share, which is what the example percentages above show.
import json
from datasets import load_dataset
ds = load_dataset("Rapidata/camera-movement", split="train")
row = ds[0]
print(row["prompt"], "|", row["model1"], "vs", row["model2"])
w1, w2 = row["weighted_results_video1_movement"], row["weighted_results_video2_movement"]
print(f"win share {row['model1']}: {w1 / (w1 + w2):.0%}")
votes = json.loads(row["detailed_results_movement"])
print(len(votes), "votes, first one:", votes[0])
This dataset combines material under different terms:
Rapidata's technology makes collecting human feedback at scale faster and more accessible than ever before. Visit rapidata.ai to learn more about how we're revolutionizing human feedback collection for AI development.
Explore our latest model rankings on our website.