Paper | Project Page | GitHub
Training data for aligning Molmo2 to perform Spatiotemporal Grounded Chain-of-Thought (SGCoT) on VET-Bench — generating explicit object tracking trajectories before answering questions.
This dataset contains 300 synthetic samples where the model generates a structured trajectory <tracks> producing a final answer. The trajectories encode spatial coordinates (x, y in a 1000×1000 normalized space) at 0.5-second intervals over 12 seconds (25 frames).
| Task | Samples | Description |
|---|---|---|
| Cup tracking | 200 | Track which cup hides a ball through a shell game |
| Card tracking | 100 | Track the Queen of Hearts through a three-card monte |
Note that you can train with any continuous trajectories, but using in-distribution trajectories causes minimal degradation of the original tracking ability.
Each sample is a chat-style messages entry in JSONL format:
{
"messages": [
{
"role": "user",
"content": [{"type": "text", "text": "Track the cup that contains the ball and answer which cup contains the ball at the end of the video."}]
},
{
"role": "assistant",
"content": [{"type": "text", "text": "<tracks coords=\"0.0 1 772 524;0.5 1 805 310;...;12.0 1 216 517\">the cup that contains the ball</tracks> Answer: left."}]
}
]
}
<tracks coords="t obj x y;t obj x y;...">object description</tracks>
left, middle, rightThe trajectories are synthetically generated from movement extracted from real tracking data:
from datasets import load_dataset
dataset = load_dataset("tiedong/Molmo2-SGCoT", split="train")
6 commits
Paper | Project Page | GitHub
Training data for aligning Molmo2 to perform Spatiotemporal Grounded Chain-of-Thought (SGCoT) on VET-Bench — generating explicit object tracking trajectories before answering questions.
This dataset contains 300 synthetic samples where the model generates a structured trajectory <tracks> producing a final answer. The trajectories encode spatial coordinates (x, y in a 1000×1000 normalized space) at 0.5-second intervals over 12 seconds (25 frames).
| Task | Samples | Description |
|---|---|---|
| Cup tracking | 200 | Track which cup hides a ball through a shell game |
| Card tracking | 100 | Track the Queen of Hearts through a three-card monte |
Note that you can train with any continuous trajectories, but using in-distribution trajectories causes minimal degradation of the original tracking ability.
Each sample is a chat-style messages entry in JSONL format:
{
"messages": [
{
"role": "user",
"content": [{"type": "text", "text": "Track the cup that contains the ball and answer which cup contains the ball at the end of the video."}]
},
{
"role": "assistant",
"content": [{"type": "text", "text": "<tracks coords=\"0.0 1 772 524;0.5 1 805 310;...;12.0 1 216 517\">the cup that contains the ball</tracks> Answer: left."}]
}
]
}
<tracks coords="t obj x y;t obj x y;...">object description</tracks>
left, middle, rightThe trajectories are synthetically generated from movement extracted from real tracking data:
from datasets import load_dataset
dataset = load_dataset("tiedong/Molmo2-SGCoT", split="train")
6 commits