
This dataset is used for the training of the LiveCC-7B-Instruct model. We only allow the use of this dataset for academic research and educational purposes. For OpenAI GPT-4o generated user prompts, we recommend users check the OpenAI Usage Policy.
After we finished the pre-training of LiveCC-7B-Base model with Live-CC-5M dataset, we trained LiveCC-7B-Instruct model with vision-language data from four sources:
Live-WhisperX-526K: This repository. It includes
2FPS Video Clips: https://huggingface.co/datasets/chenjoya/Live-WhisperX-526K/tree/main/videos
Annotation JSONL (WhisperX ASR): https://huggingface.co/datasets/chenjoya/Live-WhisperX-526K/blob/main/live_whisperx_526k_with_seeks.jsonl
It contains 527,583 real-time video commentary instances, with YouTube categories:

Each line of the JSONL file is organized in a common user/assistant conversation format with a special "text_stream" key. Example:
[
{"role": "user", "content": [{"type": "video", "video": "video/youtube/0jlPIAcUAxs.mp4", "video_start": 18.96, "video_end": 67.93}, {"type": "text", "text": "How do I replace a bicycle tire tube step by step?"}]},
{"role": "assistant", "content": [{"type": "text_stream", "text_stream": [[18.96, 19.38, "Alright,"], [19.6, 19.64, "the"], [19.64, 19.86, "first"], [19.86, 20.0, "thing"], [20.0, 20.12, "you"], [20.12, 20.24, "want"], ...]}]}
]
Each item in "text_stream" indicates start timestamp, end timestamp, and the word. Please refer to our dataloader (https://github.com/showlab/livecc/data/lmm_dataset.py) to learn how to make it compatible with popular LMMs (e.g. QwenVL series).
The last line of JSONL contains the file handle seek indices:
b'[0, 8066, 10955, 19013, 35559, 45911, 50610, 64291, 70911, 94252, ...]'
This allows for easy streaming access using:
import json
# read the last line of jsonl
def readlastline(path: str):
with open(path, "rb") as f:
f.seek(-2, 2) # avoid last \n
while f.read(1) != b"\n":
f.seek(-2, 1)
return f.readline()
# parse to seek indices list
seeks = json.loads(readlastline('live_whisperx_526k_with_seeks.jsonl'))
# during data loader
def __getitem__(self, index):
...
with open('live_whisperx_526k_with_seeks.jsonl') as f:
f.seek(seeks[index])
datum = json.loads(f.readline())
...
Use the following code to get the video path in downloaded dir:
save_video_root = 'xxx'
datum = json.loads(line)
element = datum[0]['content'][0]
file = os.path.basename(element['video'])
name, ext = os.path.splitext(file)
video_path = os.path.join(save_video_root, f"{name}_{element['video_start']:.2f}-{element['video_end']:.2f}_2.0fps{ext}")
if not os.path.exists(video_path):
video_path = video_path.replace('_2.0fps', '_2fps')


Please read the paper Section3 for details. They have been fully open-sourced at: https://github.com/showlab/livecc/tree/main/data/production
If you find our work helpful, feel free to give us a cite ;)
@article{livecc,
author = {Joya Chen and Ziyun Zeng and Yiqi Lin and Wei Li and Zejun Ma and Mike Zheng Shou},
title = {LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale},
journal = {arXiv preprint arXiv:2504.16030},
year = {2025}
}

This dataset is used for the training of the LiveCC-7B-Instruct model. We only allow the use of this dataset for academic research and educational purposes. For OpenAI GPT-4o generated user prompts, we recommend users check the OpenAI Usage Policy.
After we finished the pre-training of LiveCC-7B-Base model with Live-CC-5M dataset, we trained LiveCC-7B-Instruct model with vision-language data from four sources:
Live-WhisperX-526K: This repository. It includes
2FPS Video Clips: https://huggingface.co/datasets/chenjoya/Live-WhisperX-526K/tree/main/videos
Annotation JSONL (WhisperX ASR): https://huggingface.co/datasets/chenjoya/Live-WhisperX-526K/blob/main/live_whisperx_526k_with_seeks.jsonl
It contains 527,583 real-time video commentary instances, with YouTube categories:

Each line of the JSONL file is organized in a common user/assistant conversation format with a special "text_stream" key. Example:
[
{"role": "user", "content": [{"type": "video", "video": "video/youtube/0jlPIAcUAxs.mp4", "video_start": 18.96, "video_end": 67.93}, {"type": "text", "text": "How do I replace a bicycle tire tube step by step?"}]},
{"role": "assistant", "content": [{"type": "text_stream", "text_stream": [[18.96, 19.38, "Alright,"], [19.6, 19.64, "the"], [19.64, 19.86, "first"], [19.86, 20.0, "thing"], [20.0, 20.12, "you"], [20.12, 20.24, "want"], ...]}]}
]
Each item in "text_stream" indicates start timestamp, end timestamp, and the word. Please refer to our dataloader (https://github.com/showlab/livecc/data/lmm_dataset.py) to learn how to make it compatible with popular LMMs (e.g. QwenVL series).
The last line of JSONL contains the file handle seek indices:
b'[0, 8066, 10955, 19013, 35559, 45911, 50610, 64291, 70911, 94252, ...]'
This allows for easy streaming access using:
import json
# read the last line of jsonl
def readlastline(path: str):
with open(path, "rb") as f:
f.seek(-2, 2) # avoid last \n
while f.read(1) != b"\n":
f.seek(-2, 1)
return f.readline()
# parse to seek indices list
seeks = json.loads(readlastline('live_whisperx_526k_with_seeks.jsonl'))
# during data loader
def __getitem__(self, index):
...
with open('live_whisperx_526k_with_seeks.jsonl') as f:
f.seek(seeks[index])
datum = json.loads(f.readline())
...
Use the following code to get the video path in downloaded dir:
save_video_root = 'xxx'
datum = json.loads(line)
element = datum[0]['content'][0]
file = os.path.basename(element['video'])
name, ext = os.path.splitext(file)
video_path = os.path.join(save_video_root, f"{name}_{element['video_start']:.2f}-{element['video_end']:.2f}_2.0fps{ext}")
if not os.path.exists(video_path):
video_path = video_path.replace('_2.0fps', '_2fps')


Please read the paper Section3 for details. They have been fully open-sourced at: https://github.com/showlab/livecc/tree/main/data/production
If you find our work helpful, feel free to give us a cite ;)
@article{livecc,
author = {Joya Chen and Ziyun Zeng and Yiqi Lin and Wei Li and Zejun Ma and Mike Zheng Shou},
title = {LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale},
journal = {arXiv preprint arXiv:2504.16030},
year = {2025}
}