📄 Tech Report | 💻 GitHub
Realtime-QA-100K is a 100K-sample realtime video question answering dataset
constructed from YouTube videos. Each sample contains a multimodal
conversation and frame timestamp metadata that aligns every <|video|> token in
the assistant text with one video frame timestamp.
This repository does not contain the actual video files. We only provide annotations, YouTube video IDs, and timestamp metadata. Users are responsible for downloading the videos themselves in compliance with YouTube's Terms of Service and the original uploaders' licenses.
Figure 1. End-to-end construction pipeline of Realtime-QA-100K.
The construction of Realtime-QA-100K follows a highly structured, multi-layer data synthesis pipeline:
<|video|>, text, and <|silence|> slots.training_sample.json represented in the messages field.seed=42)<|video|> / frame_timestamps alignment failures in release build: 0The 100K release is a deterministic subsample of a larger realtime video QA
training JSONL used for streaming video LLM fine-tuning. Only rows whose video
source is YouTube are eligible; the final set is drawn uniformly at random with
seed=42.
Per-row fields are compacted for Hub distribution:
system_prompt.txt instead of
being repeated in every row.[RealTime QA][INFO] / [RealTime QA][PLACEHOLDER] markers from the
source format are removed; their information is represented by
video.video_id, video.category, video.subcategory, and
video.frame_timestamps.YouTube URL template: https://www.youtube.com/watch?v={video_id}
The dataset has one split, train, stored as Parquet shards:
data/train-00000-of-00002.parquet
data/train-00001-of-00002.parquet
Each row has the following schema:
{
"id": "rtqa_000000",
"language": "en",
"messages": [
{"role": "user", "content": ""},
{"role": "assistant", "content": "<|silence|><|video|>..."},
{"role": "user", "content": "..."},
{"role": "assistant", "content": "...<|video|>..."}
],
"video": {
"video_id": "hWYS_j90p-I",
"category": "Art",
"subcategory": "ugee_tablet",
"frame_timestamps": ["00:00", "00:01"]
}
}
Global metadata files:
system_prompt.txt: shared system prompt removed from individual rows.video_ids.txt: unique YouTube video IDs in the 100K sample.Integrity invariant. For every row, the number of <|video|> tokens appearing
in all assistant messages equals len(video.frame_timestamps). This was
checked on the full 100K build with zero mismatches.
<|video|>: marks the position corresponding to a video frame. The number of
<|video|> tokens in the assistant messages equals
len(video.frame_timestamps).<|silence|>: marks a timestep where the assistant should remain silent.<|...|>: turn-break token. Because the dataset is built for real-time
streaming, an assistant answer may still be unfolding when the video moves
on to a new scene. In that case the model needs to wrap up the current reply
quickly and yield so the next turn can respond to the new scene. <|...|>
marks exactly that intentional early cut-off — it is a designed token in the
conversation, not a text truncation or an ellipsis.pip install -U datasets pyarrow huggingface_hub
from datasets import load_dataset
ds = load_dataset("OpenMOSS-Team/Realtime-QA-100K", split="train")
print(ds)
# Dataset({features: ['id', 'language', 'messages', 'video'], num_rows: 100000})
print(ds[0])
This downloads the two Parquet shards (about 95 MB in total) into the Hugging
Face cache and returns a Dataset with four fields: id, language,
messages, and video.
from datasets import load_dataset
ds = load_dataset("OpenMOSS-Team/Realtime-QA-100K", split="train", streaming=True)
for ex in ds.take(3):
print(ex["id"], ex["language"], len(ex["video"]["frame_timestamps"]))
ds = load_dataset("OpenMOSS-Team/Realtime-QA-100K", split="train")
en_ds = ds.filter(lambda x: x["language"] == "en")
art_ds = ds.filter(lambda x: x["video"]["category"] == "Art")
print(len(en_ds), len(art_ds))
Every row omits the system message because every example uses the same prompt,
stored in system_prompt.txt. Download it once, then prepend it to each row:
from datasets import load_dataset
from huggingface_hub import hf_hub_download
ds = load_dataset("OpenMOSS-Team/Realtime-QA-100K", split="train")
system_prompt_path = hf_hub_download(
repo_id="OpenMOSS-Team/Realtime-QA-100K",
filename="system_prompt.txt",
repo_type="dataset",
)
with open(system_prompt_path, "r", encoding="utf-8") as f:
system_prompt = f.read()
def to_chat(example):
return [{"role": "system", "content": system_prompt}] + example["messages"]
chat = to_chat(ds[0])
for msg in chat[:3]:
print(msg["role"], "->", msg["content"][:80])
example = ds[0]
url = f"https://www.youtube.com/watch?v={example['video']['video_id']}"
print(url)
If you prefer raw Parquet I/O without the datasets package:
import pyarrow.parquet as pq
from huggingface_hub import hf_hub_download
shard_path = hf_hub_download(
repo_id="OpenMOSS-Team/Realtime-QA-100K",
filename="data/train-00000-of-00002.parquet",
repo_type="dataset",
)
table = pq.read_table(shard_path)
print(table.schema)
print(table.num_rows)
Or with pandas:
import pandas as pd
df = pd.read_parquet(shard_path)
print(df.head())
Video files are not included. Each example provides:
video.video_id: YouTube video ID.video.frame_timestamps: timestamps sampled at 1 fps.video.category and video.subcategory: original topic labels from data collection.Users should independently obtain videos and align frames using
video.frame_timestamps.
| Language | Samples |
|---|---|
English (en) | 93,815 |
Chinese (zh) | 6,185 |

The shared system prompt is stored separately in system_prompt.txt. Each row
therefore contains either 4 messages (one user–assistant turn) or 6 messages
(two turns).
| Messages per row | Samples |
|---|---|
| 4 | 56,720 |
| 6 | 43,280 |

len(video.frame_timestamps) per sample at 1 fps (one timestamp per second of
video coverage in the conversation).
| Metric | Value |
|---|---|
| Min | 2 |
| Median (p50) | 53 |
| Mean | 71.46 |
| Max | 240 |
The pie chart below groups lengths into 1-minute buckets; summary stats also appear in the figure legend.

| Category | Samples |
|---|---|
Technology_and_Innovation | 4,164 |
Virtual_and_Digital | 3,974 |
Cities_and_Architecture | 3,880 |
Space_And_Astronomy | 3,860 |
Military_and_Equipment | 3,853 |
Adventure_and_Extreme | 3,850 |
Children_and_Toys | 3,839 |
Weather_and_Climate | 3,837 |
Oceans_and_Waterways | 3,775 |
Animal_and_wildlife | 3,771 |
POV | 3,758 |
Nature | 3,744 |
Art | 3,744 |
Industry_and_Manufacturing | 3,736 |
Sports | 3,719 |
Daily_Life | 3,680 |
Transportation_and_Machinery | 3,659 |
Culture_and_Festivals | 3,654 |
Disasters | 3,634 |
Spectacles | 3,608 |
Geology_and_Mining | 3,595 |
Energy_and_Power | 3,594 |
Fashion_and_Design | 3,575 |
Science_and_Experiments | 3,547 |
Historical_Reenactment | 3,451 |
Medicine_and_Human_Body | 3,344 |
Agriculture_and_Horticulture | 3,155 |

The annotation data in this repository is released under CC-BY-NC-4.0.
This repository does not host or redistribute any raw video files. All video IDs and metadata belong to the respective content creators on YouTube. The copyright of the original video content remains entirely with the original owners/uploaders. This dataset only provides annotations, indices, and metadata for academic and research purposes.
Users are solely responsible for obtaining the video files independently. Any downloading or processing of YouTube videos must be done in strict compliance with YouTube's Terms of Service and the licenses selected by the original uploaders. The dataset creators are not liable for any terms of service violations or copyright infringements arising from users' independent data acquisition.
This dataset is intended for academic research on real-time and streaming video understanding. It should not be used for commercial deployment, face recognition, surveillance, or other applications that may violate privacy or platform terms.
@article{wang2026mossvideo,
title = {{MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention}},
author = {Pengyu Wang, Chenkun Tan, Shaojun Zhou, Wei Huang, Qirui Zhou, Zhan Huang, Zhen Ye, Jijun Cheng, Xiaomeng Qian, Yanxin Chen, Xingyang He, Huazheng Zeng, Chenghao Wang, Pengfei Wang, Hongkai Wang, Shanqing Gao, Yixian Tian, Chenghao Liu, Xinghao Wang, Botian Jiang, Xipeng Qiu},
year = {2026},
journal = {arXiv preprint arXiv:2606.07639},
eprint = {2606.07639},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2606.07639}
}
📄 Tech Report | 💻 GitHub
Realtime-QA-100K is a 100K-sample realtime video question answering dataset
constructed from YouTube videos. Each sample contains a multimodal
conversation and frame timestamp metadata that aligns every <|video|> token in
the assistant text with one video frame timestamp.
This repository does not contain the actual video files. We only provide annotations, YouTube video IDs, and timestamp metadata. Users are responsible for downloading the videos themselves in compliance with YouTube's Terms of Service and the original uploaders' licenses.
Figure 1. End-to-end construction pipeline of Realtime-QA-100K.
The construction of Realtime-QA-100K follows a highly structured, multi-layer data synthesis pipeline:
<|video|>, text, and <|silence|> slots.training_sample.json represented in the messages field.seed=42)<|video|> / frame_timestamps alignment failures in release build: 0The 100K release is a deterministic subsample of a larger realtime video QA
training JSONL used for streaming video LLM fine-tuning. Only rows whose video
source is YouTube are eligible; the final set is drawn uniformly at random with
seed=42.
Per-row fields are compacted for Hub distribution:
system_prompt.txt instead of
being repeated in every row.[RealTime QA][INFO] / [RealTime QA][PLACEHOLDER] markers from the
source format are removed; their information is represented by
video.video_id, video.category, video.subcategory, and
video.frame_timestamps.YouTube URL template: https://www.youtube.com/watch?v={video_id}
The dataset has one split, train, stored as Parquet shards:
data/train-00000-of-00002.parquet
data/train-00001-of-00002.parquet
Each row has the following schema:
{
"id": "rtqa_000000",
"language": "en",
"messages": [
{"role": "user", "content": ""},
{"role": "assistant", "content": "<|silence|><|video|>..."},
{"role": "user", "content": "..."},
{"role": "assistant", "content": "...<|video|>..."}
],
"video": {
"video_id": "hWYS_j90p-I",
"category": "Art",
"subcategory": "ugee_tablet",
"frame_timestamps": ["00:00", "00:01"]
}
}
Global metadata files:
system_prompt.txt: shared system prompt removed from individual rows.video_ids.txt: unique YouTube video IDs in the 100K sample.Integrity invariant. For every row, the number of <|video|> tokens appearing
in all assistant messages equals len(video.frame_timestamps). This was
checked on the full 100K build with zero mismatches.
<|video|>: marks the position corresponding to a video frame. The number of
<|video|> tokens in the assistant messages equals
len(video.frame_timestamps).<|silence|>: marks a timestep where the assistant should remain silent.<|...|>: turn-break token. Because the dataset is built for real-time
streaming, an assistant answer may still be unfolding when the video moves
on to a new scene. In that case the model needs to wrap up the current reply
quickly and yield so the next turn can respond to the new scene. <|...|>
marks exactly that intentional early cut-off — it is a designed token in the
conversation, not a text truncation or an ellipsis.pip install -U datasets pyarrow huggingface_hub
from datasets import load_dataset
ds = load_dataset("OpenMOSS-Team/Realtime-QA-100K", split="train")
print(ds)
# Dataset({features: ['id', 'language', 'messages', 'video'], num_rows: 100000})
print(ds[0])
This downloads the two Parquet shards (about 95 MB in total) into the Hugging
Face cache and returns a Dataset with four fields: id, language,
messages, and video.
from datasets import load_dataset
ds = load_dataset("OpenMOSS-Team/Realtime-QA-100K", split="train", streaming=True)
for ex in ds.take(3):
print(ex["id"], ex["language"], len(ex["video"]["frame_timestamps"]))
ds = load_dataset("OpenMOSS-Team/Realtime-QA-100K", split="train")
en_ds = ds.filter(lambda x: x["language"] == "en")
art_ds = ds.filter(lambda x: x["video"]["category"] == "Art")
print(len(en_ds), len(art_ds))
Every row omits the system message because every example uses the same prompt,
stored in system_prompt.txt. Download it once, then prepend it to each row:
from datasets import load_dataset
from huggingface_hub import hf_hub_download
ds = load_dataset("OpenMOSS-Team/Realtime-QA-100K", split="train")
system_prompt_path = hf_hub_download(
repo_id="OpenMOSS-Team/Realtime-QA-100K",
filename="system_prompt.txt",
repo_type="dataset",
)
with open(system_prompt_path, "r", encoding="utf-8") as f:
system_prompt = f.read()
def to_chat(example):
return [{"role": "system", "content": system_prompt}] + example["messages"]
chat = to_chat(ds[0])
for msg in chat[:3]:
print(msg["role"], "->", msg["content"][:80])
example = ds[0]
url = f"https://www.youtube.com/watch?v={example['video']['video_id']}"
print(url)
If you prefer raw Parquet I/O without the datasets package:
import pyarrow.parquet as pq
from huggingface_hub import hf_hub_download
shard_path = hf_hub_download(
repo_id="OpenMOSS-Team/Realtime-QA-100K",
filename="data/train-00000-of-00002.parquet",
repo_type="dataset",
)
table = pq.read_table(shard_path)
print(table.schema)
print(table.num_rows)
Or with pandas:
import pandas as pd
df = pd.read_parquet(shard_path)
print(df.head())
Video files are not included. Each example provides:
video.video_id: YouTube video ID.video.frame_timestamps: timestamps sampled at 1 fps.video.category and video.subcategory: original topic labels from data collection.Users should independently obtain videos and align frames using
video.frame_timestamps.
| Language | Samples |
|---|---|
English (en) | 93,815 |
Chinese (zh) | 6,185 |

The shared system prompt is stored separately in system_prompt.txt. Each row
therefore contains either 4 messages (one user–assistant turn) or 6 messages
(two turns).
| Messages per row | Samples |
|---|---|
| 4 | 56,720 |
| 6 | 43,280 |

len(video.frame_timestamps) per sample at 1 fps (one timestamp per second of
video coverage in the conversation).
| Metric | Value |
|---|---|
| Min | 2 |
| Median (p50) | 53 |
| Mean | 71.46 |
| Max | 240 |
The pie chart below groups lengths into 1-minute buckets; summary stats also appear in the figure legend.

| Category | Samples |
|---|---|
Technology_and_Innovation | 4,164 |
Virtual_and_Digital | 3,974 |
Cities_and_Architecture | 3,880 |
Space_And_Astronomy | 3,860 |
Military_and_Equipment | 3,853 |
Adventure_and_Extreme | 3,850 |
Children_and_Toys | 3,839 |
Weather_and_Climate | 3,837 |
Oceans_and_Waterways | 3,775 |
Animal_and_wildlife | 3,771 |
POV | 3,758 |
Nature | 3,744 |
Art | 3,744 |
Industry_and_Manufacturing | 3,736 |
Sports | 3,719 |
Daily_Life | 3,680 |
Transportation_and_Machinery | 3,659 |
Culture_and_Festivals | 3,654 |
Disasters | 3,634 |
Spectacles | 3,608 |
Geology_and_Mining | 3,595 |
Energy_and_Power | 3,594 |
Fashion_and_Design | 3,575 |
Science_and_Experiments | 3,547 |
Historical_Reenactment | 3,451 |
Medicine_and_Human_Body | 3,344 |
Agriculture_and_Horticulture | 3,155 |

The annotation data in this repository is released under CC-BY-NC-4.0.
This repository does not host or redistribute any raw video files. All video IDs and metadata belong to the respective content creators on YouTube. The copyright of the original video content remains entirely with the original owners/uploaders. This dataset only provides annotations, indices, and metadata for academic and research purposes.
Users are solely responsible for obtaining the video files independently. Any downloading or processing of YouTube videos must be done in strict compliance with YouTube's Terms of Service and the licenses selected by the original uploaders. The dataset creators are not liable for any terms of service violations or copyright infringements arising from users' independent data acquisition.
This dataset is intended for academic research on real-time and streaming video understanding. It should not be used for commercial deployment, face recognition, surveillance, or other applications that may violate privacy or platform terms.
@article{wang2026mossvideo,
title = {{MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention}},
author = {Pengyu Wang, Chenkun Tan, Shaojun Zhou, Wei Huang, Qirui Zhou, Zhan Huang, Zhen Ye, Jijun Cheng, Xiaomeng Qian, Yanxin Chen, Xingyang He, Huazheng Zeng, Chenghao Wang, Pengfei Wang, Hongkai Wang, Shanqing Gao, Yixian Tian, Chenghao Liu, Xinghao Wang, Botian Jiang, Xipeng Qiu},
year = {2026},
journal = {arXiv preprint arXiv:2606.07639},
eprint = {2606.07639},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2606.07639}
}